跳到正文
原文
Jonas Geiping· @jonasgeiping · X·本站收录 · 原文发表

CIAware 基准更新:GPT-6 Astra 几乎饱和识别 AI 控制干预

AI 导读

Jonas Geiping 更新了 CIAware 基准,该基准衡量前沿模型是否能察觉施加在其文本上的 AI 控制干预。他引用 Joachim Schaeffer 的内容指出,GPT-6 Astra 在这项任务上表现极强,接近饱和;而在今年 5 月发布初版基准时,多数模型仅略高于随机水平,当时还有读者质疑该基准不切实际。Schaeffer 认为这是一个大问题,会挑战控制协议的有效性:模型对控制干预的高识别度意味着干预会泄露监控器的信息,从而有助于绕过监控。

正文

We just updated CI-aware bench, which measures whether frontier models are aware of AI control interventions made to their text.

When we first talked about this in spring, most models were near chance, and some readers complained that the bench was unrealistic...

Astra nearly saturates it now.

For more details, check out Joachim's thread linked below:

引用Joachim Schaeffer@JSchaeff3r
Frontier models can now detect AI control interventions. In particular, GPT-6 Astra is incredibly good at this. When we published the initial version of our CIAware benchmark in May most models were not much above chance level, now a step change happened. I think this is a big problem and challenges the effectiveness of control protocols. High control intervention awareness means that interventions leaks information about the monitor which helps circumventing the monitor. Details in the thread 🧵
在 X 查看被引用的帖子

来源:Jonas Geiping · x.com