Anthropic对齐压力测试团队谈失准模型研究
AI 导读
@EvanHub 领导 Anthropic 的对齐压力测试团队,该团队有两项职责:作为"第二道防线"审查自身的安全工作,以及构建失准的"模式生物"来研究模型可能如何欺骗性地行事。他在我们 2024 年湾区对齐研讨会上的演讲: https://youtu.be/JfDlbzF6rsY
正文
@EvanHub leads Anthropic's Alignment Stress Testing team, which has two jobs: acting as a "second line of defense" reviewing its own safety work, and building "model organisms" of misalignment to study how models might behave deceptively. His talk from our 2024 Bay Area Alignment Workshop:
https://youtu.be/JfDlbzF6rsY