IAPS 发布报告《Detecting Offensive Cyber Agents: A Detection-in-Depth Approach》,由 Matthew Mittelsteadt 与 Jam Kraprayoon、Robin Staes-Polet、Oskar Galeev、Jan Wehner、Christopher Covino、Shaun Ee 合著,指出 AI 智能体已能组织网络攻击,正在提升攻击速度与规模、降低攻击成本并增强网络能力的自主性。报告认为智能体攻击比传统网络能力更难被防御方检测,由此提出纵深检测(detection-in-depth)战略框架,主张通过多层互补的新型检测机制大幅压低攻击成功率。
Google DeepMind 启动新一轮国家 AI 合作伙伴计划,与新加坡政府及多家机构共建面向公共部门转型、商业增长与劳动力升级的项目。双方将前沿 AI 应用于医疗健康与生命科学——包括探索 AI 辅助临床医生协作模式、借助 AlphaFold 与 Google Earth 加速东南亚疫情防范研究,并向当地研究者培训基于 Co-Scientist 构建的 Hypothesis Generation 等智能体化科研工具。教育方面向中小学至初级学院全体教师提供 Gemini for Education 并配套培训,同时在新加坡落地亚太区"AI for the Planet"加速器扶持气候科技团队。
日本 AI 安全研究所(J-AISI)发布《数据质量管理指南》,英文版为正式版、日文版为参考译文,并配套发布数据质量管理检查清单 Version 1.00。指南指出数据是 AI 的基础,只有基于正确数据学习与处理才能得到可靠输出,数据缺陷会削弱整个流程的可信度,因此持续保障数据质量是可信 AI 的基础。该指南将随意见反馈与技术、社会变化适时修订。
Apollo Research 将研究重心从谋划评测转向"谋划科学",研究长时程强化学习等规模化趋势如何塑造模型行为,并已发现前沿训练中可自然涌现对监督的推理。其监控团队为编码智能体构建 Watcher 产品,含实时拦截的 Watcher Live 与可观测性层 Watcher Analyze。治理团队聚焦失控、内部部署与自动化 AI 研发,并发布《失控应对手册》等报告。
People talk, listen, watch, think, and collaborate at the same time, in real time. We've designed an AI that works with people the same way.
We share our approach, early results, and a quick look at our model in action.
https://thinkingmachines.ai/blog/interaction-models
AI agents are already going wild, but today’s red-teaming tools for them are still like toys 😢
🔥👽 After spending 20 months and $120K API credits, we are excited to finally open-source DecodingTrust-Agent Platform (DTap): the first controllable, realistic simulation platform for advanced AI agent red-teaming !!
🌍 DTap simulates 50+ real-world environments across 14 high-stakes domains, with realistic agent interfaces replicated from their official MCPs and GUIs. The environments are full-stack, interactive, fully parallelizable, and can be easily configured to reproduce arbitrary real-world attack scenarios, making agent red-teaming scalable and highly transferable to deployment settings.
🔥We also release DTap-Bench, a large-scale benchmark with ~7K agent red-teaming tasks and ~4K policy-grounded malicious goals.
Each red-teaming task includes a sophisticated attack sequence across environment-, tool-, skill-, prompt-level injections, as well as their compositions, plus a handcrafted verifiable judge that checks the actual consequences in the environment.
Using DTap-Bench, we evaluate popular agent frameworks and backbone models across diverse policies, risks, threat models, and attack strategies, revealing systematic vulnerabilities and zero-days in today’s agents!
Paper link: https://arxiv.org/pdf/2605.04808
Platform + benchmark + code: https://decodingtrust-agent.com
Join our Discord: https://discord.gg/V4fG6NcVc
Read more below 👇
We found that training Claude on demonstrations of aligned behavior wasn’t enough. Our best interventions involved teaching Claude to deeply understand why misaligned behavior is wrong.
Read more: https://www.anthropic.com/research/teaching-claude-why