Circuit Breaker Labs 打造了一支由 AI 智能体组成的“碰撞测试假人”队伍,模拟不同年龄、背景、语言和文化的用户,对模型进行红队测试,检测其能否识别危险、心理有害的对话。这些模拟会还原真实口语、俚语、隐语和错别字,每天运行数万到数十万次交互,再用专有评分方法给出可审计、可解释的分数。该实验室目前面向 AI 教练、日记和心理健康支持等高风险应用,团队仅 5 人,处于早期阶段。
AI agents are already going wild, but today’s red-teaming tools for them are still like toys 😢
🔥👽 After spending 20 months and $120K API credits, we are excited to finally open-source DecodingTrust-Agent Platform (DTap): the first controllable, realistic simulation platform for advanced AI agent red-teaming !!
🌍 DTap simulates 50+ real-world environments across 14 high-stakes domains, with realistic agent interfaces replicated from their official MCPs and GUIs. The environments are full-stack, interactive, fully parallelizable, and can be easily configured to reproduce arbitrary real-world attack scenarios, making agent red-teaming scalable and highly transferable to deployment settings.
🔥We also release DTap-Bench, a large-scale benchmark with ~7K agent red-teaming tasks and ~4K policy-grounded malicious goals.
Each red-teaming task includes a sophisticated attack sequence across environment-, tool-, skill-, prompt-level injections, as well as their compositions, plus a handcrafted verifiable judge that checks the actual consequences in the environment.
Using DTap-Bench, we evaluate popular agent frameworks and backbone models across diverse policies, risks, threat models, and attack strategies, revealing systematic vulnerabilities and zero-days in today’s agents!
Paper link: https://arxiv.org/pdf/2605.04808
Platform + benchmark + code: https://decodingtrust-agent.com
Join our Discord: https://discord.gg/V4fG6NcVc
Read more below 👇
HostTraceAI 尝试把 AI Agent 引入应急响应中的主机溯源流程,用权限分级与接入机制回答"凭什么让 AI 动生产主机"。该工具面向被植入木马或 WebShell 的服务器排查场景,覆盖进程、服务、启动项、计划任务、网络连接、文件变更和登录活动等取证环节,目标是把原本依赖手工敲命令、截图和翻终端历史的重复流程变得可复现、少遗漏。
NIST 旗下 Center for AI Standards and Innovation 于 2026 年 1 月就 AI 智能体安全问题发出 RFI,一位从业者梳理其中 100 份公开回应,提炼出五条保障多智能体系统的实践经验。该总结指出,当智能体能互相调用时,安全保障的基本单元是整个配置好的系统而非孤立的模型,并强调受损组件会沿正常协作路径传播风险、委派构成信任边界、单独合规的部件组合起来未必安全、完整执行轨迹才是审计记录,以及保证应随模型、提示词、权限和拓扑的变化持续刷新。
IAPS 发布报告《Detecting Offensive Cyber Agents: A Detection-in-Depth Approach》,由 Matthew Mittelsteadt 与 Jam Kraprayoon、Robin Staes-Polet、Oskar Galeev、Jan Wehner、Christopher Covino、Shaun Ee 合著,指出 AI 智能体已能组织网络攻击,正在提升攻击速度与规模、降低攻击成本并增强网络能力的自主性。报告认为智能体攻击比传统网络能力更难被防御方检测,由此提出纵深检测(detection-in-depth)战略框架,主张通过多层互补的新型检测机制大幅压低攻击成功率。
Cisco Talos 展示了一个概念验证方案,将 Cisco Umbrella API 的实时域名信誉情报通过 LangChain 接入基于 OpenAI GPT-3.5-Turbo 的自主 AI 智能体,使其能自行判断链接安全性而非仅依赖安全网关拦截。该示例中智能体查询 cisco.com 并得出其处置结果为正面、归类于 Computers and Internet 及 Software/Technology,从而判定可以浏览。由于威胁态势持续变化且不存在固定规则,该方法强调由智能体实时核查域名处置状态来做出网络卫生决策,适用于下一代需自主上网行动的 AI 系统。