跳到正文
原文
Transformer Circuits(Anthropic 可解释性)·· 2026-04-02精选关注度62

Anthropic 在 Claude Sonnet 4.5 中发现情绪概念表征并证实其影响输出

Emotion Concepts and their Function in a Large Language Model

AI 导读

Anthropic 的可解释性研究团队在 Claude Sonnet 4.5 中发现了情绪概念的内部表征,并表明这些表征会因果性地影响模型的输出。研究以 Transformer Circuits 文章形式发布,但当前可见内容仅为摘要,未提供具体探测方法、表征位置和实验细节。

推荐理由

Anthropic 在 Claude Sonnet 4.5 中找到情绪概念表征并给出其因果影响输出的证据,是机制可解释性研究的一个具体案例。

阅读原文transformer-circuits.pub