UK AISI 开源复现 Anthropic 奖励作弊引发失准研究
AI 导读
UK AISI 模型透明团队用开源模型、RL 环境、算法与工具链,复现了 Anthropic 的《Natural Emergent Misalignment from Reward Hacking in Production RL》研究,并分享了一个与思维链忠实性相关的意外结果。该复现由 @satvikgolechha 以 7 条推文线程发布,Thomas Read 转发称这是其团队的新工作。
正文
more exciting work from my team at UK AISI! an open source reproduction of "Natural Emergent Misalignment from Reward Hacking in Production RL"
Research from Model Transparency @ UK AISI: we reproduce the Anthropic work "Natural Emergent Misalignment from Reward Hacking in Production RL" using OS models, RL environments, algorithms, and tooling + we share an unexpected result related to CoT faithfulness.
🧵 (1 of 7)在 X 查看被引用的帖子