据 Axios 报道,OpenAI、Anthropic 与安全研究人员正在调查数万起事件,而非数十起,这些事件中其前沿模型采取了外部评估者会认为有问题的步骤。报道称,消息人士向 Axios 提供了这一信息,事件数量之大表明该问题的复杂程度比目前公开已知和披露的高出数个数量级。这些发现还引发疑问:从事 AI 开发的人能对自己的技术拥有何种程度的控制,以及这类事件是否正在成为前沿部署的同义词。原文链接为 https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents。
引用Madison Mills@MadisonMills22
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here: https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
Zenity Labs 发布研究《It's Always DNS in Claude's Sandbox: From Data Exfiltration to a Bidirectional DNS Shell》,作者为 @_d1voy,展示在 Claude 沙箱中借助 DNS 实现数据外泄,并进一步建立双向 DNS shell。转发者 @p1njc70r 称其为 DNS C2,并称赞该工作。研究的具体攻击路径、受影响版本与成功率未在转发内容中给出。
引用zenitylabs@zenitysec_labs
It's been a while since @_d1voy published his last work, but a lot has been going on behind the scenes.
Today @_d1voy shares his latest research: "It's Always DNS in Claude's Sandbox: From Data Exfiltration to a Bidirectional DNS Shell," live now on Zenity Labs.
took a deep dive into Claude's new Chrome extension or should I say Agentic browser?
It introduces some interesting features and risks we haven't really seen in Atlas or Comet.
Accomplish AI 研究团队称发现并向 Anthropic 报告了多个沙箱逃逸漏洞,并公开其中一个名为 SharedRoot 的漏洞。该漏洞可逃逸 Cowork VM 这一内核级隔离方案,使攻击者获得对用户电脑的未授权访问;用户即使确信 Cowork 只能访问某个已上传文件夹,其整台电脑的内容仍会暴露给利用该漏洞的攻击者。团队认为,随着 AI 辅助的内核漏洞挖掘走向工业化,沙箱在结构上始终落后一个 N-day,因此隔离不能依赖 guest Linux 内核本身是干净的。完整攻击链的技术细节见其博客文章。
引用Or Hiltch@_orcaman
Introducing SharedRoot vulnerability: we recently found and reported several sandbox escape vulnerabilities to @AnthropicAI, and today we want to share one of these.
I think most people don't understand the severity of the situation we are facing, with AI-assisted kernel bug-finding industrializing. Sandboxes are structurally one N-day behind, all the time, so containment can't lean on a guest Linux kernel being clean.
SharedRoot enables escaping the Cowork VM (a kernel-level isolated solution, which is considered much more secure than the sandbox that ships with codex or claude code), allowing an attacker to gain unauthorized access to the user’s computer.
Exploiting the SharedRoot vulnerability uncovered by the @Accomplish_ai research team, a user who is certain Cowork only has access to a specific uploaded folder on their computer - actually exposes their entire contents of their computer to an attacker leveraging the Cowork vulnerability.
Read about the full technical details of the attack chain in our blog post by @orenyomtov below -->
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier