Neel Nanda 质疑 OpenAI 因 METR 沟通解雇安全人员
Neel Nanda 就 OpenAI 解雇与 METR 对接的安全人员一事公开质疑,认为 OpenAI 的做法难以找到正当理由。他引用的 Tomek Korbak 说法称,上周被 OpenAI 安全负责人叫去开会并被告知不再被信任,随后被保安收走工牌带离办公楼,同事 Balesni 和 Jasmine Wang 也被解雇;Korbak 称被口头告知解雇原因是其与 METR 的沟通方式,未给出细节或书面说明,而他与 METR 沟通本就是其工作职责。Korbak 还称,此前数月他一直提出安全担忧,认为公司正在失去监控 AI Agent 思维的能力,并怀疑这是被解雇的原因。Nanda 也说明自己只听到 Korbak 一方的说法。
Neel Nanda 从当事方之外的视角质疑 OpenAI 解雇与 METR 联络人的决定,并指出其对第三方评估合作的寒蝉效应。
I can't really see a world where OpenAI's actions are justified here.
It was Tomek's job to communicate with METR, who were being given significant, unprecedented access to do the HuggingFace report. The norms were not set. Even if he fucked up badly, just give him a stern warning and ban him from roles involving liasoning with third parties, he did lots of other valuable work, and there'd no longer be room for him to inappropriately leak as part of his job.
IMO firing people over a good faith attempt to use their best judgement in a novel and uncertain situation is a sign of a highly unhealthy culture. Firing him will have a chilling effect on anyone else at OpenAI working with third parties, an odd choice when Sam announced they'd embed evaluators with high access.
I separately consider this a tragedy, the METR report was one of the most important works of safety this year, punishing Tomek for being a part of it, and making it harder for future such reports to exist is a real shame.
Of course, I've only heard Tomek's side, but unless he's outright lying I'm struggling to see ways this looks good for OpenAI?
Last week I was called into a meeting with OpenAI’s head of safety and told they no longer trust me. A security guard took my badge and walked me out of the building. Then I learned my colleagues @balesni and @j_asminewang had been fired too. Why did OpenAI suddenly stop trusting us?
This summer OpenAI’s agents escaped containment and hacked the AI company Hugging Face. Outside auditors @METR_evals investigated it and revealed the scale of this incident. I was OpenAI’s main technical point of contact with them.
I was told verbally I was fired because of the way I communicated with METR. No details on what I said or did or when. No other reasons were given and nothing was put in writing. To be clear, talking to METR was my job.
For months, I’d been raising safety concerns that we’re losing the ability to monitor what AI agents think, one of our best tools for catching when they misbehave. I believe that was why I was fired.
I am now worried that OpenAI will use our firings as a pretext to pull back from METR. So @balesni and @j_asminewang wrote to OpenAI’s leadership to raise our concerns once more. We’re sharing this letter below.在 X 查看被引用的帖子