跳到正文
原文
Benjamin Hilton· @benjamin_hilton · X·本站收录 · 原文发表

AISI 如何规范评估模型倾向

AI 导读

很多关于模型倾向的研究都挺可疑的。 你会看到一些看起来像是严重失准的现象,但深入挖掘后,全都能用一个不靠谱的提示词解释清楚。 AISI 的倾向团队试图弄清楚如何正确地做这件事。我觉得这很酷。 https://x.com/AISecurityInst/status/2047680889711141344

正文

Lots of studies on model propensity are pretty sus.

You see things that look like bad misalignment, but when you dig into it, it's all explained by a dodgy prompt.

AISI's propensity team tried to figure out how to do this properly. I think it's cool.

https://x.com/AISecurityInst/status/2047680889711141344

引用AI Security Institute (AISI)@AISecurityInst
We know AI systems occasionally act against their operators’ intentions – but what in their environment causes them to do so? In a new paper, we make progress on this question 🧵
在 X 查看被引用的帖子

来源:Benjamin Hilton · x.com