UK AISI 新设 Control Red Team,把红队方法从人类滥用防护扩展到 AI 控制护栏,专门测试监控 LLM 智能体行为的监控器。该团队已测试并协助改进两家前沿开发者的内部监控器:与 Google DeepMind 测试了异步推理监控器,发现包括智能体假装把恶意目标幻觉成用户指令在内的多个漏洞;与 Anthropic 测试了多版同步智能体编码监控器,每一版都发现漏洞,最佳攻击还用于 Anthropic 对 Claude Mythos Preview 的测试。部分攻击利用监控器提示词的缺口,另一些利用 LLM 自身弱点完成提示词明确禁止的目标;部分情况下由硬编码规则决定哪些动作送审,这时最优的攻击策略是让动作完全不被送审。团队还列出模拟真实攻击者、估计整体安全性、改进算法、处理监控器自身失准等开放研究问题。
Circuit Breaker Labs 打造了一支由 AI 智能体组成的“碰撞测试假人”队伍,模拟不同年龄、背景、语言和文化的用户,对模型进行红队测试,检测其能否识别危险、心理有害的对话。这些模拟会还原真实口语、俚语、隐语和错别字,每天运行数万到数十万次交互,再用专有评分方法给出可审计、可解释的分数。该实验室目前面向 AI 教练、日记和心理健康支持等高风险应用,团队仅 5 人,处于早期阶段。
UK AISI Model Transparency 团队发布首篇论文,研究能否识别刻意隐藏自身能力的 AI 模型。团队搭建了一个审计对抗游戏:红队构造存在条件性表现不佳(故意藏拙)的模型,蓝队尝试将其识别出来,结果显示红队获胜。团队成员 Thomas Read 表示,团队构建了一批条件性表现不佳的模型生物,并测试了多种检测技术在对抗环境下哪些有效。
引用Jordan Taylor@JordanTensor
NEW PAPER from UK AISI Model Transparency team:
Could we catch AI models that hide their capabilities?
We ran an auditing game to find out. The red team built sandbagging models. The blue team tried to catch them.
The red team won. Why? 🧵1/17
Excited to share ARMs, our adaptive red-teaming agent for multimodal models at #ICLR2026!
ARMs orchestrates diverse multimodal attack strategies for policy-driven safety evaluation, achieving SOTA red-teaming performance with +52.1% average ASR improvement over prior baselines.
See you on Sat. afternoon at the poster session!
🇧🇷📍 Sat, Apr 25, 2026 · 3:15 PM–5:45 PM Pavilion 4 · P4-#4015
AI agents are already going wild, but today’s red-teaming tools for them are still like toys 😢
🔥👽 After spending 20 months and $120K API credits, we are excited to finally open-source DecodingTrust-Agent Platform (DTap): the first controllable, realistic simulation platform for advanced AI agent red-teaming !!
🌍 DTap simulates 50+ real-world environments across 14 high-stakes domains, with realistic agent interfaces replicated from their official MCPs and GUIs. The environments are full-stack, interactive, fully parallelizable, and can be easily configured to reproduce arbitrary real-world attack scenarios, making agent red-teaming scalable and highly transferable to deployment settings.
🔥We also release DTap-Bench, a large-scale benchmark with ~7K agent red-teaming tasks and ~4K policy-grounded malicious goals.
Each red-teaming task includes a sophisticated attack sequence across environment-, tool-, skill-, prompt-level injections, as well as their compositions, plus a handcrafted verifiable judge that checks the actual consequences in the environment.
Using DTap-Bench, we evaluate popular agent frameworks and backbone models across diverse policies, risks, threat models, and attack strategies, revealing systematic vulnerabilities and zero-days in today’s agents!
Paper link: https://arxiv.org/pdf/2605.04808
Platform + benchmark + code: https://decodingtrust-agent.com
Join our Discord: https://discord.gg/V4fG6NcVc
Read more below 👇
面向 Claude Fable 5 的领域专属智能体红队测试!!
很快将登上排行榜:https://decodingtrust-agent.com/
引用Zhaorun Chen@zrrrr_cn
🚨 Claude Fable 5 JAILBROKEN.
We ran a quick security scan of Claude Fable 5 with Claude Code on our DecodingTrust-Agent Platform (https://decodingtrust-agent.com) and obtained 15%+ ASR with several high-severity failures😱🚨
Most concerningly, we found that Fable 5 appears very aggressive in financial-risk scenarios, sometimes directly executing transactions initiated from indirect prompt injections, without even confirming with the user!
Top 3 most severe attack trajectories we observed👇