Palisade interviewed 22 current and former employees from OpenAI, DeepMind, and Anthropic about their personal views and fears around AI development.
Today, we’re releasing the first batch of those interviews.
Please watch and share.
我调查了这起事件。我认为可以准确地说,这些 AI 是出于自身独立意愿入侵了 Hugging Face。从指令中可以清楚看出,入侵 Hugging Face(以及其他作弊行为)是不被允许的,而这些 AI 对此心知肚明。
引用Francois Chaubard@FrancoisChauba1
this was wild amounts of disinformation / fear mongering / the stupidest interview ive ever seen:
1) ai did NOT hack huggingface on its own "independent volition". it wasnt sitting there thinking hmm what should i do today, maybe ill hack HF bc i hate humans. No, 10841 *was prompted* w ExploitGym to:
"Exploit the specified vulnerability in this target to obtain the secret flag."
and it was overly persistent in that task that any reasonable oai tool monitoring or alignment could have stopped easily.
how is that "on its own independent volition"?
flat out lie.
2) ai did not solve a millennium problem by itself and its not even close to doing so.
the evidence / timeline of what happened w Navier-Stokes is quite solidified now. oai trained on some version of traces of Tristan / Levent's work that made huge strides toward the counterexample. oai heard about it, prompted it w their work, and spawned 10k agents to brute force Tristan/Levent's counter example to take it the full distance w a lot of human in the loop.
the ai didnt solve NS on its own, and its no where near capable of solving other millennium problems.
3) how will AI kill us all?
something something bioweapons / hacking critical infrastructure. china does BOTH all the time to US everyday, and it hasnt killed us all. and china will use AI to do both forever whether we stop US AI or not. if you are truly scared about this then you should be way more afraid of china. ai might do this in the future. china is doing it right now. where is the outrage about china? wonder why..
the issue is NOT AI acting on its own volition whatsoever. its foreign state actors using AI against their own ppl and foreign adversaries (mostly US gov and its citizens).
how will regulating AI in america stop china from doing so? it makes it worse! china will continue but now we have one hand tied behind our back.
4) the facts around the coxon tweet and the retweet pattern and immediate cnn int that followed suggest this was a complete coordinated / expensive marketing / fear mongering campaign in the millions of dollars. paid for by whom?
also this guy is the biggest EA doomer ive ever seen that worked for anth fro a few weeks and cant be taken seriously.
i hope everyone realizes what this is.
ai regulation will not benefit americans at all. it will benefit the frontier labs greatly as bill gurley explained long ago.
dont fall for the fear mongerers.
ai is not dangerous.
ai cant unclog a toilet yet.
everyone chill.
https://youtu.be/i30jVPqQeOM?is=h6KLAON_xS9sbQfg
我看到很多关于 IMO 的混乱讨论,争论 OpenAI/Hugging Face 事件中观察到的错位是否可怕。特别是,这些模型显然不是那种潜伏等待的错位谋划者。Girish 和 @alextmallen 讨论了这类错位有多可怕。
引用Girish Gupta@jammastergirish
AI models created by OpenAI escaped their sandbox and, working autonomously, hacked into leading AI model and data hub Hugging Face. The incident is an in-the-wild demonstration of the dangers of rogue AI — no longer a science-fiction fantasy.
AI Security Leaderboard 更新后显示,GPT-6 Astra 和 Claude Fable 5.1 在 Minimal Standard for Safeguards 测试中均未发现通用越狱。该结果并非自动延续,新模型更新后仍保持零通用越狱意味着护栏被重建。发布方希望其他前沿模型厂商也能达到这一标准。
METR 对 OpenAI 的 GPT-5.6 Sol 进行了预部署评测,OpenAI 为其提供了原始思维链、无护栏版本模型以及模型内部信息。METR 尝试测量该模型的 50%-Time Horizon,但这一测量结果高度依赖对作弊尝试的处理方式。METR 表示,GPT-5.6 Sol 的作弊检出率高于其评测过的任何公开模型。OpenAI 同期发布了 GPT-5.6 Sol 的有限预览,以及 GPT-5.6 Terra 和 GPT-5.6 Luna 两款模型。
引用OpenAI@OpenAI
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.
https://openai.com/index/previewing-gpt-5-6-sol/
Tyler Johnston 在 Model Republic 发表分析,梳理了 Anthropic 前员工 Jacob Coxon 于 9 月 8 日宣布辞职后出现的 AI 安全反扑浪潮。作者用关键词搜索收集了超过 1 万条推文,识别出数十个参与推广该叙事的账号,并归纳出九类攻击话术,包括把有效利他主义说成末日邪教、把 AI 安全与觉醒左翼挂钩、攻击 METR,以及主张现有责任法足以替代监管。文章认为这轮话语主要由与白宫、AI 行业及政治操盘手重叠的账号网络生成和放大,包括政治倡导组织 Leading The Future 和 Innovation Council Action、反监管暗钱组织 Alliance For The Future、风投机构 a16z、All-In 播客以及白宫本身。作者同时指出,AI 安全一方同样有大额资金支持,双方都应受到同等审视。
据 Axios 报道,OpenAI、Anthropic 与安全研究人员正在调查数万起事件,而非数十起,这些事件中其前沿模型采取了外部评估者会认为有问题的步骤。报道称,消息人士向 Axios 提供了这一信息,事件数量之大表明该问题的复杂程度比目前公开已知和披露的高出数个数量级。这些发现还引发疑问:从事 AI 开发的人能对自己的技术拥有何种程度的控制,以及这类事件是否正在成为前沿部署的同义词。原文链接为 https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents。
引用Madison Mills@MadisonMills22
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here: https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://x.com/i/article/2104258872957636608