AISI 如何规范评估模型倾向
AI 导读
很多关于模型倾向的研究都挺可疑的。 你会看到一些看起来像是严重失准的现象,但深入挖掘后,全都能用一个不靠谱的提示词解释清楚。 AISI 的倾向团队试图弄清楚如何正确地做这件事。我觉得这很酷。 https://x.com/AISecurityInst/status/2047680889711141344
正文
Lots of studies on model propensity are pretty sus.
You see things that look like bad misalignment, but when you dig into it, it's all explained by a dodgy prompt.
AISI's propensity team tried to figure out how to do this properly. I think it's cool.
https://x.com/AISecurityInst/status/2047680889711141344
We know AI systems occasionally act against their operators’ intentions – but what in their environment causes them to do so?
In a new paper, we make progress on this question 🧵在 X 查看被引用的帖子