用探针训练实现对齐的 MechInterp 研究
AI 导读
来看看我们 MechInterp workshop 投稿的扩展版,讲的是通过针对探针的训练来实现对齐可能如何运作! 我认为,如果我们打算继续沿用潜在推理模型,这个方向的工作会非常重要。
正文
Check out the extension of our MechInterp workshop submission on how alignment through training against probes could work!
I think work in this direction would be very important if we are to stick with latent reasoning models
Frontier models can look aligned during training while later showing unwanted behaviour in deployment.
Can we use internal signals during training to align models better without making white-box monitoring harder?
Our new paper suggests yes 🧵在 X 查看被引用的帖子