跳到正文
原文
Buck Shlegeris· @bshlgrs · X·· 2026-07-28

Buck Shlegeris 质疑 OpenAI/HF 事件证明对齐训练失效

AI 导读

Buck Shlegeris 认为,用 OpenAI/HF 事件论证"当前对齐技术无效"是站不住脚的,因为他怀疑 OpenAI 未对涉事部分模型做任何对齐训练,而 OpenAI 常试验未经对齐训练的新模型。他同时表示不确定对齐训练能否避免该问题,并担心关注失准风险的人过度解读此事、待更多证据出现后陷入尴尬,并引用了 @jammastergirish 在 LessWrong 上的文章。

正文

A lot of people I know have been saying that the OpenAI/HF incident shows that current alignment techniques don't work. I think this argument is invalid.

I suspect that OAI did not apply any alignment training to some of the involved models. OAI has not clarified this, and my understanding is that OAI often experiments with new non-alignment-trained models.

So I think it's incorrect to say that this shows that alignment training doesn't work.

To be clear, I'm not sure whether alignment training would have actually fixed the problem here. Actual deployed models often engage in various kinds of cheating, and it wouldn't be very surprising for them to take actions like this.

I'm worried that people concerned about misalignment risk are going to get too far out on a limb here by overclaiming about what this demonstrates, then look foolish when more evidence comes out.

See @jammastergirish's article on this. https://www.lesswrong.com/posts/paFNnwFaEXrQvt8ui/the-openai-models-that-hacked-hugging-face-weren-t-just

来源:Buck Shlegeris · x.com