前沿模型已能识别 AI 控制干预,GPT-6 Astra 表现突出
AI 导读
研究者 Joachim Schaeffer 表示,前沿模型如今已能识别 AI 控制干预,其中 GPT-6 Astra 在这方面的能力尤为突出。他所在的 CIAware 基准在 5 月发布初版时,多数模型的识别能力仅略高于随机水平,现在则出现了阶跃式变化。作者认为这会削弱控制协议的有效性:模型对控制干预的高度感知会泄露监控方的信息,从而有助于绕过监控。
正文
Frontier models can now detect AI control interventions.
In particular, GPT-6 Astra is incredibly good at this. When we published the initial version of our CIAware benchmark in May most models were not much above chance level, now a step change happened.
I think this is a big problem and challenges the effectiveness of control protocols. High control intervention awareness means that interventions leaks information about the monitor which helps circumventing the monitor.
Details in the thread 🧵