METR 的 Ryan Greenblatt 呼吁 AI 公司公开架构可监控性权衡证据
METR 的 Ryan Greenblatt 对 AI 架构转向以不透明激活而非思维链进行推理(即"neuralese"架构)表示担忧,认为 Astra 是这一方向上令人不安的一步。他指出公开信息不足以就 Astra 架构与训练方法改动在可监控性与性能之间的权衡展开充分讨论,呼吁 AI 公司发布相关证据并公开其政策,Redwood AI 也提出了追踪无 CoT 推理能力与可监控性政策的提案。他强调公司应谨慎对待可能消除或大幅削弱对思维链依赖的架构。
I'm very worried about changes to AI architectures that result in AIs thinking in opaque activations instead of in chain of thought (aka "neuralese" architectures).
Based on limited public evidence, it seems like Astra was a concerning step in this direction.
Unfortunately, there was insufficient public information to have a well-informed public scientific discussion about whether the changes to architectures and training methods that went into Astra were a good trade-off between monitorability and performance. We also don't know how AI companies will make these trade-offs going forward. I think companies should release the evidence needed for a reasonably informed public conversation about how these trade-offs should be made. They should also make and publish their policies around this. We've written up a proposal for how this could work.
I don't know if this proposal will be sufficient to avoid the most concerning architectures, but it seems like a relatively robust step in the right direction. AI companies should be very cautious about pursuing architectures that could eliminate or greatly reduce dependence on chain of thought, and certainly shouldn't do this before the rest of the world has a chance to discuss the evidence and their policies.
Some architectures could weaken CoT monitorability, or remove the CoT altogether.
We've written a proposal for how companies could be transparent about no-CoT reasoning abilities, other monitorability evidence, and policies for preserving monitorability. https://www.redwoodresearch.org/blog/proposal-for-tracking-architecture-on-monitorability在 X 查看被引用的帖子