研究者对多模态大语言模型的“thinking-with-images”范式做因果审计,发现视觉工具调用带来的准确率提升并非普遍来自返回的视觉证据。作者把视觉工具使用建模为因果图,区分观察中介路径与动作引发的捷径,并在策略、轨迹、步骤三个层面做干预,其中步骤层面的 Visual Evidence Gain 用于分离单次返回观察的贡献。在六个代表性模型和五个细粒度感知基准上,审计发现两类失败模式:Calling Without Looking 中返回的观察对答案没有因果影响,Looking Without Planning 中观察有信息但调用时序不连贯。轨迹层面的诊断显示,策略层面的准确率增益集中在少数校准良好的样本上,作者将这种总体增益与因果无效并存的现象称为视觉工具使用的假象。代码已公开,论文被 EMNLP 2026 Findings 收录。
Jonas Geiping 等人发布预印本,提出直接针对无害性与诚实性探针优化模型,并称在持续更新探针的前提下效果良好,模型学会对有害请求生成无害回答、在压力下保持诚实。作者转述的论文观点认为,AI 安全领域不少被视为禁忌的技术(如用思维链检测奖励作弊、用模型内部表征做训练)缺乏清晰科学依据,而随着未来模型可能靠通用奖励寻求行为刷满对齐训练场景,用内部表征监督训练且不丧失可监控性将愈发重要。作者本人指出,这种训练让模型学会无害,而不是学会拒答有害请求,小模型在被诱导输出有害内容时会出现一些有意思的回答。
So the Assistant adopts traits more from human characters who it resembles. We exploit this to learn about *how* the model represents the Assistant. E.g. the model treats the Assistant as resembling elite-school humans more than non-elite ones.
(Is this because the model trusts elite-school people more in determining what to believe? We think not because papers like Slocum et al 2025 suggest that provenance doesn't matter for belief uptake from finetuning.)
@NeelNanda5 is widely regarded as one of the top two experts on mechanistic interpretability in the world.
“Speaking as an interpretability expert, please do not rely on us to save you on the current trajectory."
https://youtu.be/J38ot52b2-E
Are SAEs dead?
Will they save us from neuralese?
Should I just use a probe?
We get these questions all the time. Part 2 of our educational series on applied interpretability explains what SAEs are good for, when *not* to use them, and what to use instead. 🧵
METR 的 Ryan Greenblatt 对 AI 架构转向以不透明激活而非思维链进行推理(即"neuralese"架构)表示担忧,认为 Astra 是这一方向上令人不安的一步。他指出公开信息不足以就 Astra 架构与训练方法改动在可监控性与性能之间的权衡展开充分讨论,呼吁 AI 公司发布相关证据并公开其政策,Redwood AI 也提出了追踪无 CoT 推理能力与可监控性政策的提案。他强调公司应谨慎对待可能消除或大幅削弱对思维链依赖的架构。
引用Redwood Research@redwood_ai
Some architectures could weaken CoT monitorability, or remove the CoT altogether.
We've written a proposal for how companies could be transparent about no-CoT reasoning abilities, other monitorability evidence, and policies for preserving monitorability. https://www.redwoodresearch.org/blog/proposal-for-tracking-architecture-on-monitorability