跳到正文
原文
Thomas Read· @thjread · X·本站收录 · 原文发表

UK AISI 复现 Anthropic 的 steering 方法抑制评测感知

AI 导读

UK AISI Model Transparency 团队复现了 Anthropic 用于抑制评测感知(evaluation awareness)的 steering 方法。最意外的发现是,作为对照的 steering 向量(内容与书架上的书有关)产生的效果与刻意设计的向量一样大。该结果提示在评测感知相关实验中,对照向量的选择可能显著影响结论。

正文

New from the UK AISI Model Transparency team: we replicated Anthropic's steering approach for suppressing evaluation awareness. Our most surprising finding: "control" steering vectors (about books on shelves!) can have effects as large as deliberately designed ones. 🧵

来源:Thomas Read · x.com