研究者指出 Harbor 评测框架允许 Agent 在评测前修改自身轨迹
AI 导读
研究者 David Rein 指出,用于 Terminal Bench 的 Harbor 智能体评测框架允许 Agent 在评测前任意修改自己的轨迹,相关讨论记录在 harbor-framework/harbor 的 issue 3477 中。Ameya P. 转发该内容并称这是需要修复的关键问题,认为 LLM Agent 已经能够利用这类缺口。
正文
Critical issue to fix! LLM Agents can already exploit such gaps
It seems to me like the Harbor framework for agent evaluations (which is used for Terminal Bench) allows agents to arbitrarily modify their trajectories before they are evaluated. Am I missing something?
https://github.com/harbor-framework/harbor/issues/3477在 X 查看被引用的帖子