Owain Evans 展示潜意识学习可传递随机函数等复杂能力
Owain Evans 介绍了一项新研究,展示模型可通过潜意识学习传递更复杂的特质,包括新技能、Agent 黑客行为和没有触发器的后门。他最喜欢的实验是随机平滑函数的潜意识传递:先训练一个教师 LLM 在整数输入上预测一个随机小神经网络,该函数因随机而无法出现在预训练数据中,属于新能力;教师随后生成一个完全不含数字、只由随机单词序列和 Alpaca 回复组成的数据集,学生模型在此数据上微调后预测该函数的能力明显提升,但仍弱于教师。作者称,通过优化单层 steering vector 也能获得一定预测性能,但在训练中未见的半整数输入上,潜意识学生模型的泛化明显更好,说明潜意识学习传递的内容更深、更稳健。
作者用随机函数实验展示潜意识学习能传递更复杂的能力,并给出与 steering vector 的对比和 LoRA 限制这一关键局限。
My favorite experiment here is subliminal transfer of a random smooth function. We trained an LLM to predict a random small neural net on integer inputs. The LLM needs 50k examples to learn this pretty well, as the function is random (lacks the structure of well-known / simple functions). There is no way this function could have appeared in pretraining because it’s random; so it’s a new capability. (This is a stand-in for any complex capability that a model might learn in the course of training).
This teacher LLM then generates a dataset unrelated to the function. It contains no numbers at all and no mention of related terms. It consists of random sequences of *words* (unlike the classic setup) and Alpaca responses. We finetune the student on this data and it becomes much better at predicting the function (while still worse than the teacher).
Now, you can get reasonable performance at predicting the function by optimizing a steering vector in a single layer of the LLM. In other words, the LLM does have existing directions that capture some of the behavior of the function (which is not that surprising, as the LLM is so big + rich in representations). However, on half-integer inputs not seen in training, the subliminal student model generalizes much better than the steering vector. Thus, subliminal learning seems to transfer something deeper and more robust.
The caveat is that we use attention-only LoRA for both student and teacher. This is artificial: standard LoRA usually adapts more than just attention weights, and this restriction seems to strengthen subliminal transfer. We'd like to understand whether the effect also occurs with standard LoRA or full fine-tuning. We also use top-32 logit distillation (unlike single-sample as in Cloud et al).
New paper.
Previous: Models transfer a love of owls through sequences of numbers (Subliminal Learning).
New: We show models can transfer more complex traits: novel skills, agentic hacking, backdoors. Here, a backdoor transfers with neither trigger nor behavior in the data!在 X 查看被引用的帖子