Apollo:谋划安全论证的四大主张
AI 导读
要安全地开发前沿 AI,开发者必须能够证明其模型没有在谋划,即没有在追求非预期目标时暗中与之作对。 新文章:任何谋划安全论证都必须做出的四项主张,以及嵌入式评估者验证这些主张所需的资源 🧵
正文
To develop frontier AI safely, developers must be able to show that their models are not scheming, i.e. covertly working against them in pursuit of unintended goals.
New post: four claims any scheming safety case must make, and the resources embedded evaluators need to verify them 🧵