跳到正文
原文
Alexander Panfilov· @kotekjedi_ml · X·· 3 小时前

用探针训练实现对齐的 MechInterp 研究

AI 导读

来看看我们 MechInterp workshop 投稿的扩展版,讲的是通过针对探针的训练来实现对齐可能如何运作! 我认为,如果我们打算继续沿用潜在推理模型,这个方向的工作会非常重要。

正文

Check out the extension of our MechInterp workshop submission on how alignment through training against probes could work!

I think work in this direction would be very important if we are to stick with latent reasoning models

引用Lena Libon @ COLM 2026@lenalibon
Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal signals during training to align models better without making white-box monitoring harder? Our new paper suggests yes 🧵
在 X 查看被引用的帖子

来源:Alexander Panfilov · x.com