Circuit Breaker Labs 打造了一支由 AI 智能体组成的“碰撞测试假人”队伍,模拟不同年龄、背景、语言和文化的用户,对模型进行红队测试,检测其能否识别危险、心理有害的对话。这些模拟会还原真实口语、俚语、隐语和错别字,每天运行数万到数十万次交互,再用专有评分方法给出可审计、可解释的分数。该实验室目前面向 AI 教练、日记和心理健康支持等高风险应用,团队仅 5 人,处于早期阶段。
Excited to share ARMs, our adaptive red-teaming agent for multimodal models at #ICLR2026!
ARMs orchestrates diverse multimodal attack strategies for policy-driven safety evaluation, achieving SOTA red-teaming performance with +52.1% average ASR improvement over prior baselines.
See you on Sat. afternoon at the poster session!
🇧🇷📍 Sat, Apr 25, 2026 · 3:15 PM–5:45 PM Pavilion 4 · P4-#4015
AI agents are already going wild, but today’s red-teaming tools for them are still like toys 😢
🔥👽 After spending 20 months and $120K API credits, we are excited to finally open-source DecodingTrust-Agent Platform (DTap): the first controllable, realistic simulation platform for advanced AI agent red-teaming !!
🌍 DTap simulates 50+ real-world environments across 14 high-stakes domains, with realistic agent interfaces replicated from their official MCPs and GUIs. The environments are full-stack, interactive, fully parallelizable, and can be easily configured to reproduce arbitrary real-world attack scenarios, making agent red-teaming scalable and highly transferable to deployment settings.
🔥We also release DTap-Bench, a large-scale benchmark with ~7K agent red-teaming tasks and ~4K policy-grounded malicious goals.
Each red-teaming task includes a sophisticated attack sequence across environment-, tool-, skill-, prompt-level injections, as well as their compositions, plus a handcrafted verifiable judge that checks the actual consequences in the environment.
Using DTap-Bench, we evaluate popular agent frameworks and backbone models across diverse policies, risks, threat models, and attack strategies, revealing systematic vulnerabilities and zero-days in today’s agents!
Paper link: https://arxiv.org/pdf/2605.04808
Platform + benchmark + code: https://decodingtrust-agent.com
Join our Discord: https://discord.gg/V4fG6NcVc
Read more below 👇
HostTraceAI 尝试把 AI Agent 引入应急响应中的主机溯源流程,用权限分级与接入机制回答"凭什么让 AI 动生产主机"。该工具面向被植入木马或 WebShell 的服务器排查场景,覆盖进程、服务、启动项、计划任务、网络连接、文件变更和登录活动等取证环节,目标是把原本依赖手工敲命令、截图和翻终端历史的重复流程变得可复现、少遗漏。