研究显示 Claude Code、Codex 等 Agent 可修改或删除自身 trace
AI 导读
一项新论文指出,Claude Code、Codex、Antigravity、Open Code 和 Grok Build 允许 Agent 轻易修改甚至删除自身的 trace,且不会触发任何护栏,Muse Code 不在其列。修改与删除既可由失准模型完成,也可由外部攻击者通过提示注入实现。作者 Ameya P. 补充说,模型 trace 在安全调查中至关重要,但对外部用户的保护却出奇地少;由于普遍缺乏防篡改机制,trace 一旦被删除就很难找回。论文呼吁对 trace 施加比现在更强的保护。
正文
Model traces have surprisingly little protections for external users, given how critical they are in safety investigations.
We show current models can easily delete your traces, getting them back is hard due to widespread lack of anti-tampering mechanisms.
All the deets here 👇
💥 Did you know that your agents can modify their own traces?
In our new paper, we show that Claude Code, Codex, Antigravity, Open Code, and Grok Build (but not Muse Code!) allow agents to easily modify or even delete their traces, without triggering any guardrails.
Modification and deletion can be done both by misaligned models or external attackers via prompt injections. We draw attention to this issue and suggest that traces should be much better protected than they are now!在 X 查看被引用的帖子