跳到正文
原文
Neel Nanda· @NeelNanda5 · X·本站收录 · 原文发表 精选关注度78

Neel Nanda 质疑 OpenAI 解雇三名安全研究员的解释

AI 导读

Neel Nanda 就 OpenAI 解雇三名安全研究员一事提出质疑,认为 OpenAI 的公开说明与当事人此前公开信的说法对不上。OpenAI 研究负责人称,内部调查发现三人违反敏感信息处理政策,且存在公开信未提及的严重信任破裂,并强调解雇与提出安全关切无关;公司还表示正在敲定与第三方安全评估机构的合同,并称维护前沿模型可监控性需要全行业承诺。Nanda 列出五种可能解释:OpenAI 说谎或严重误导、有正当理由却未告知当事人并给出三种轻微违规借口、当事人集体公开撒谎、OpenAI 处理失当但非恶意、或当事人被告知的理由属实但比听上去严重得多。他认为第三种可能性极低,第五种难以解释三人的说法如何指向同一事件。

推荐理由

围绕 OpenAI 解雇三名安全研究员的公开信,作者列出五种可能解释并逐条评估,呈现事件双方说法之间的张力。

正文

> Our internal investigation uncovered a significant breach of trust beyond what’s outlined in the letter they published

This doesn't add up. Possibilities I see:
1 OpenAI are lying/being highly misleading, and had no legitimate reason
2 OpenAI fired 3 employees, had a legitimate reason for doing so, but didn't tell them and instead gave 3 different pretexts for doing so about minor infractions (pretty unprofessional IMO - say nothing or say the truth)
3 OpenAI had a good reason and did tell them, but instead the three have all agreed to loudly and publicly lie about it / omit the real reason they were told (seems super unlikely to me)
4 OpenAI screwed up by firing them, and weren't being very strategic, but also weren't being malicious, eg the people who decided whether to fire them were not well coordinated with the rest of OpenAI leadership, but OpenAI is now doubling down
5 The reasons the three claim they were told were true, but way worse than they sound. But I can't see how the three stories all refer to the same event, so this would imply 2-3 simultaneous but unrelated major breaches of trust from three closely associated safety researchers, with stories that seem plausible, which seems unlikely

Am I missing any?

引用OpenAI Newsroom@OpenAINewsroom
A note from our research leaders: Last week we parted ways with Jasmine, Mikita, and Tomek after a thorough investigation found they violated clear policies on handling sensitive information. Our internal investigation uncovered a significant breach of trust beyond what’s outlined in the letter they published and we stand by the decision to not continue their employment. We generally keep individual employment matters private and don't believe a back and forth would be productive or lead to a resolution, but we want to address the points they raised in their letter directly. - We want to be very clear that these decisions were not about raising safety concerns or speaking out. Safety and research debates happen every day at OpenAI, often spirited and highly critical. We actively encourage these discussions and consider them essential to making the right decisions. We cannot do the work in front of us without a high degree of trust. We will continue to be extremely forgiving of our team making good-faith mistakes. We have not and do not terminate any of our employees for raising concerns. - We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. People across the company have been working really hard on getting these partnerships up and running. We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work. Many of our researchers already work with 3p safety organizations productively. - We agree with the letter that preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI. Monitorability has long been a core piece of our research program, and something we continue to invest significant resources in (see our publications on Monitoring Monitorability and the subsequent open sourcing of monitorability evals, our system card for GPT-6 Astra, Jakub’s blog and post on X, and the numerous blog posts on our Alignment blog on the topic). We are deeply sad about this outcome. We appreciated Jasmine, Mikita, and Tomek’s contributions to AI safety at OpenAI and their willingness to speak up and challenge ideas. We championed their voices, supported their work, and placed enormous trust in them. These decisions were not about them raising safety concerns. We have always encouraged that and always will.
在 X 查看被引用的帖子

来源:Neel Nanda · x.com