跳到正文
原文
Jonas Geiping· @jonasgeiping · X·本站收录 · 原文发表

研究者用探针训练对齐模型,让模型学会无害而非拒答

AI 导读

Jonas Geiping 等人发布预印本,提出直接针对无害性与诚实性探针优化模型,并称在持续更新探针的前提下效果良好,模型学会对有害请求生成无害回答、在压力下保持诚实。作者转述的论文观点认为,AI 安全领域不少被视为禁忌的技术(如用思维链检测奖励作弊、用模型内部表征做训练)缺乏清晰科学依据,而随着未来模型可能靠通用奖励寻求行为刷满对齐训练场景,用内部表征监督训练且不丧失可监控性将愈发重要。作者本人指出,这种训练让模型学会无害,而不是学会拒答有害请求,小模型在被诱导输出有害内容时会出现一些有意思的回答。

正文

We recently posted our preprint about aligning models through training against probes, see below!

What I found especially interesting is that we can train models to 'be harmless', instead of training them to refuse harmful requests (and still be potentially harmful if refusal is avoided).

For smaller models, this leads to some interesting answers when we try to get harmful answers out of this harmless model (that doesn't quite know what refusals are yet)

引用Maksym Andriushchenko@maksym_andr
💥 New paper: AI safety is full of "forbidden techniques" (using CoT to detect reward hacking, using model internals for training, etc). But do they really have a clear scientific basis? I'm not sure. In this paper, we directly optimize models against harmlessness and honesty probes, and it works just fine (if you continuously update the probe!). The models learn how to generate harmless responses to harmful queries and honest responses under pressure to lie. Figuring out how to correctly use interp techniques for training is becoming increasingly important: it's very likely that soon we won't be able to align models using output-based supervision. Future models will just max out all alignment training scenarios, but for the wrong reasons due to their general reward-seeking behavior. To have a chance of aligning future models, we need to do much more research on supervising model training using their internals *without losing monitorability*!
在 X 查看被引用的帖子

来源:Jonas Geiping · x.com