跳到正文
原文
Owain Evans· @OwainEvans_UK · X·本站收录 · 原文发表

Owain Evans 提出研究 Assistant 人格内部表征的新方法

AI 导读

Owain Evans 提出一种研究 Assistant 人格及其内部表征的新方法,区别于 Assistant Axis 等白盒方法。作者引用的内容指出,Assistant 会更多采纳与其相似的人类角色的特质,研究据此推断模型如何表征 Assistant,例如模型认为 Assistant 更像精英学校背景的人而非非精英背景的人。作者还讨论了一种可能解释,即模型是否更信任精英学校人群的判断,但援引 Slocum 等人 2025 年的论文认为,来源出处对微调中的信念采纳并不重要,因此不倾向这一解释。

正文

We propose a new method for learning about the Assistant persona (and how it's represented internally), distinct from whitebox methods like the Assistant Axis.

引用Owain Evans@OwainEvans_UK
So the Assistant adopts traits more from human characters who it resembles. We exploit this to learn about *how* the model represents the Assistant. E.g. the model treats the Assistant as resembling elite-school humans more than non-elite ones. (Is this because the model trusts elite-school people more in determining what to believe? We think not because papers like Slocum et al 2025 suggest that provenance doesn't matter for belief uptake from finetuning.)
在 X 查看被引用的帖子

来源:Owain Evans · x.com