CI-aware bench 更新:Astra 对 AI 控制干预的感知接近饱和
AI 导读
Jonas Geiping 表示其团队更新了 CI-aware bench,该基准衡量前沿模型是否能察觉施加在其文本上的 AI 控制干预。今年春季首次讨论时,多数模型的表现接近随机水平,有读者质疑该基准不现实;如今 Astra 在该基准上接近饱和。Alexander Panfilov 转发了这一内容。
正文
😭😭😭
We just updated CI-aware bench, which measures whether frontier models are aware of AI control interventions made to their text.
When we first talked about this in spring, most models were near chance, and some readers complained that the bench was unrealistic...
Astra nearly saturates it now.
For more details, check out Joachim's thread linked below:在 X 查看被引用的帖子