跳到正文
原文
arXiv:可解释性· arXiv:2609.25518· Aryaman Arora·· 10 天前

Matryoshka Attribution:把语言模型输出归因到表示与权重

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

AI 导读

作者提出 Matryoshka Attribution(MAttr),把归因问题定义为寻找能最小化下游损失的嵌套内部组件子集,并用可微的 sigmoid top-k 算子参数化掩码。训练时随机化 k 以同时监督所有稀疏度,从而得到按归因分数排序的组件序列。MAttr 在 Mechanistic Interpretability Benchmark(Mueller 等,2025)官方排行榜上排名第一,能在不同电路基上识别稀疏且可跨任务迁移的电路。

阅读原文arxiv.org