CI-aware bench 更新:Astra 对 AI 控制干预的感知接近饱和
AI 导读
CI-aware bench 更新后,前沿模型对施加在其文本上的 AI 控制干预的感知能力大幅上升,Astra 在该基准上接近饱和。该基准衡量模型是否意识到自身文本被施加了 AI 控制干预。据 @jonasgeiping 介绍,今年春季首次讨论该基准时,多数模型的表现接近随机水平,当时还有读者质疑基准不切实际。Joachim Schaeffer 表示,模型在控制干预基准上如此之快取得令人担忧的分数,出乎他们的意料。
正文
We were really surprised to see models getting worrying scores on control intervention benchmark so quickly...
We just updated CI-aware bench, which measures whether frontier models are aware of AI control interventions made to their text.
When we first talked about this in spring, most models were near chance, and some readers complained that the bench was unrealistic...
Astra nearly saturates it now.
For more details, check out Joachim's thread linked below:在 X 查看被引用的帖子