研究者提出 Context-Fractured Decomposition 攻击,工具型 LLM Agent 越狱成功率最高提升 28.3 个百分点
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps
AI 导读
研究者提出 Context-Fractured Decomposition(CFD)攻击,针对工具型 LLM Agent 的“溯源缺口”部署失效模式。该攻击保留早期交互中看似无害的中间工件,在后续不同 Agent 实例或工作流阶段通过单独无害的工具动作触发有害行为,风险只在延迟的工件中介组合下显现。作者用 trace 级诊断刻画这一失效模式,并提出溯源血缘标记(provenance lineage tagging)作为可验证的缓解方向。在 Agent 系统越狱基准上,CFD 相比当前最优基线成功率最高提升 28.3 个百分点,且能绕过较强的单轮判定器。