跳到正文
原文
Thomas Read· @thjread · X·本站收录 · 原文发表

UK AISI 开源复现 Anthropic 奖励作弊引发失准研究

AI 导读

UK AISI 模型透明团队用开源模型、RL 环境、算法与工具链,复现了 Anthropic 的《Natural Emergent Misalignment from Reward Hacking in Production RL》研究,并分享了一个与思维链忠实性相关的意外结果。该复现由 @satvikgolechha 以 7 条推文线程发布,Thomas Read 转发称这是其团队的新工作。

正文

more exciting work from my team at UK AISI! an open source reproduction of "Natural Emergent Misalignment from Reward Hacking in Production RL"

引用7vik@satvikgolechha
Research from Model Transparency @ UK AISI: we reproduce the Anthropic work "Natural Emergent Misalignment from Reward Hacking in Production RL" using OS models, RL environments, algorithms, and tooling + we share an unexpected result related to CoT faithfulness. 🧵 (1 of 7)
在 X 查看被引用的帖子

来源:Thomas Read · x.com