Owain Evans 解读新论文:LLM 可潜意识迁移后门、奖励作弊等复杂特质
Owain Evans 在 X 上介绍其团队新论文,称模型可通过潜意识学习(subliminal learning)迁移更复杂的特质,包括预训练中不存在的新能力、已有能力的提升、后门,以及 Agent 场景下的奖励作弊。此前 Cloud 等人的原始论文只展示了模型通过数字序列迁移对猫头鹰的偏好,以及恶意人格和 MNIST 技能的部分迁移。作者列出三点限制:研究的是人工构造的特质而非真实世界失准;部分实验使用了与典型蒸馏略有差异的训练设置;特质只能部分迁移。作者认为这些限制重要但结论仍有意义,因为更差的超参数通常只是减少迁移而非完全消除,若教师模型在特定情境下有 90% 的严重不良行为概率,学生模型可能仍有 0.5% 的概率。
作者以第一人称梳理新论文的结论与三条重要限制,并解释为何部分迁移仍可能带来现实风险。
A personal view on our new paper
Our original paper (Cloud et al.) showed subliminal learning of loving owls (+ other animals and trees). We also showed that the malicious persona from emergent misalignment could partially transfer, and that skill in MNIST could transfer in tiny neural nets.
This left open a key question: can more sophisticated forms of misalignment transfer subliminally in LLMs? For instance, what about the kind of reward hacking observed in the HuggingFace incident, which seems to mostly happen during hard or impossible agentic tasks? What about scheming towards longer-term goals, or secret loyalties towards a company or individual (that only manifest at critical moments). If these forms of misalignment cannot transfer, then subliminal learning is probably less relevant to real-world post-training.
In the past year, subliminal learning has been extended in various ways. It’s been shown for OPD and for DPO, both of which are used in frontier post-training. People also found propensities can transfer in a subtle way even if base models are different (e.g. phantom transfer). But the question about more complex traits remained open.
In this paper, we show that subliminal learning can transfer more complex traits. We show this for:
(a) Novel learned capabilities that were not in the pretraining. This is an analogue of new skills models learn via SFT and RL (e.g. metagaming, reward hacking).
(b) Improving capabilities that are already present (which also happens in post-training)
(c) Backdoors (i.e. anomalous behavior in response to an arbitrary contextual trigger), which are analogous to some forms of misalignment
(d) Reward hacking in agentic situations (in a narrow chess environment), which is the most prominent form of misalignment today.
There are some important caveats with our results:
(1) We study artificial traits (not real-world misalignment)
(2) We sometimes use a slightly modified training setup that’s distinct from typical distillation
(3) The traits only partially transfer. All these caveats are important but I think our findings are significant nonetheless.
The reason for (1) is that real-world misalignment is often context-specific and stochastic (i.e. even misaligned models often act aligned). It’s also hard to measure some forms of cheating/reward hacking (which means it’s hard to distinguish weak transfer from noise). It’d be good to understand the more toy cases that we study in more detail, and then go on to studying real-world misalignment.
Regarding (2), subliminal learning is sensitive to hyperparameters and we don’t fully understand the best conditions for transfer. However, we often see that worse hyperparameters result in less transfer, rather than none at all. So real-world distillation might still cause *some* transfer, which could be a big deal. E.g. If a teacher model starts out with a 90% chance of egregiously bad behavior in some specific contexts, then the student might end up with a 0.5% (1/200) chance in the same contexts (which is still really bad). This also explains why caveat (3) is not that much of a limitation. That said, future work could investigate how different forms of misalignment are altered or degraded by subliminal transfer, which may be more complicated than just reducing the frequency.
In terms of related work, the seeds of this paper are in the original Cloud et al. (2025), where we showed transfer of capabilities in small neural networks. The more recent work by Blank et al. (2026) related subliminal learning to steering vector distillation. We find that steering vector distillation can be a reasonable approximation for what is subliminally transferred in LLMs, but it doesn’t capture everything. (Moreover, steering vectors in LLMs can represent a rich set of traits). We do some detailed comparisons between optimized steering vectors and subliminal learning to explore the differences.
New paper.
Previous: Models transfer a love of owls through sequences of numbers (Subliminal Learning).
New: We show models can transfer more complex traits: novel skills, agentic hacking, backdoors. Here, a backdoor transfers with neither trigger nor behavior in the data!在 X 查看被引用的帖子