Owain Evans 提出研究 Assistant 人格内部表征的新方法
AI 导读
Owain Evans 提出一种研究 Assistant 人格及其内部表征的新方法,区别于 Assistant Axis 等白盒方法。作者引用的内容指出,Assistant 会更多采纳与其相似的人类角色的特质,研究据此推断模型如何表征 Assistant,例如模型认为 Assistant 更像精英学校背景的人而非非精英背景的人。作者还讨论了一种可能解释,即模型是否更信任精英学校人群的判断,但援引 Slocum 等人 2025 年的论文认为,来源出处对微调中的信念采纳并不重要,因此不倾向这一解释。
正文
We propose a new method for learning about the Assistant persona (and how it's represented internally), distinct from whitebox methods like the Assistant Axis.
So the Assistant adopts traits more from human characters who it resembles. We exploit this to learn about *how* the model represents the Assistant. E.g. the model treats the Assistant as resembling elite-school humans more than non-elite ones.
(Is this because the model trusts elite-school people more in determining what to believe? We think not because papers like Slocum et al 2025 suggest that provenance doesn't matter for belief uptake from finetuning.)在 X 查看被引用的帖子