2025年模型对齐度显著提升
AI 导读
有趣的趋势:2025年全年,模型的对齐程度大幅提升。 自动审计发现的失准行为比例一直在下降,不仅是在Anthropic,在GDM和OpenAI也是如此。
正文
Interesting trend: models have been getting a lot more aligned over the course of 2025.
The fraction of misaligned behavior found by automated auditing has been going down not just at Anthropic but for GDM and OpenAI as well.