AI 安全排行榜更新:GPT-6 Astra 与 Claude Fable 5.1 零通用越狱
AI 导读
AI Security Leaderboard 更新后显示,GPT-6 Astra 和 Claude Fable 5.1 在 Minimal Standard for Safeguards 测试中均未发现通用越狱。该结果并非自动延续,新模型更新后仍保持零通用越狱意味着护栏被重建。发布方希望其他前沿模型厂商也能达到这一标准。
正文
The AI Security Leaderboard is updated with the latest model releases from OpenAI and Anthropic.
We are glad to report that we found zero universal jailbreaks in GPT-6 Astra and zero in Claude Fable 5.1, both tested against our Minimal Standard for Safeguards.
Robustness at this level does not carry over from one release to the next automatically. Holding at zero through a new model update means the safeguards were rebuilt. That is the standard we hope to see other frontier model providers meet.
Explore the leaderboard. Link in thread. 👇