跳到正文
原文
Goodfire Research·· 2025-12-05

Goodfire:对抗样本不是缺陷,而是叠加表示

Adversarial Examples Are Not Bugs, They Are Superposition - Goodfire

AI 导读

Goodfire 研究指出对抗样本并非模型缺陷,而是叠加表示(superposition)的产物。该研究来自 Goodfire Research,其 Silico 可解释性智能体用于解释、调试并精确控制模型行为。

阅读原文goodfire.com