跳到正文
原文
arXiv:Agent 与 MCP 安全· arXiv:2609.16098· Xiaoyan Li·· 17 天前

研究者提出面向工具集成 LLM Agent 的通用对抗攻击防御方法

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

AI 导读

研究者提出一套统一框架下的通用防御策略,用于应对工具集成 LLM Agent 面临的直接提示注入、间接提示注入、记忆投毒和后门攻击四类对抗攻击。方法包含两类基于工具的防御:Attacker Tool Filtering 用异常检测(如 Isolation Forest)识别并移除可疑工具,Normal Tool Recalling 则在规划前以白盒方式恢复 Agent 的原始工具集;同时引入 Chain-of-Thought 提示与自我反思等基于提示的防御,并通过任务改写缓解攻击。

阅读原文arxiv.org