跳到正文
原文
Ameya P.· @AmyPrb · X·本站收录 · 原文发表

研究者指出 Harbor 评测框架允许 Agent 在评测前修改自身轨迹

AI 导读

研究者 David Rein 指出,用于 Terminal Bench 的 Harbor 智能体评测框架允许 Agent 在评测前任意修改自己的轨迹,相关讨论记录在 harbor-framework/harbor 的 issue 3477 中。Ameya P. 转发该内容并称这是需要修复的关键问题,认为 LLM Agent 已经能够利用这类缺口。

正文

Critical issue to fix! LLM Agents can already exploit such gaps

引用david rein@idavidrein
It seems to me like the Harbor framework for agent evaluations (which is used for Terminal Bench) allows agents to arbitrarily modify their trajectories before they are evaluated. Am I missing something? https://github.com/harbor-framework/harbor/issues/3477
在 X 查看被引用的帖子

来源:Ameya P. · x.com