Ryan Greenblatt 加入 METR 调查 AI 风险
Ryan Greenblatt 宣布加入 METR,继续开展类似其 Hugging Face 报告那样的调查工作。他认为当前关于 AI 开发的大量基础信息未公开,而近期事件让他改变了对公开信息价值的怀疑态度;获取 AI 公司内部经核实的信息尤为紧迫,因为有限公开证据与“即将到来的递归自我改进可能大幅加速能力进展、甚至在一半年内产生极端超人类通用能力”的可能性相符。METR 初期将聚焦能力/起飞、对齐与控制,他希望其他团队覆盖安全、内部流程等领域。
I'm joining METR to work on more investigations like our Hugging Face report.
Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind.
Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term.
Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.)
While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.