Hendrycks:AI 有意识也不该交出控制权
即便 AI 产生了意识,也不意味着人类应该把控制权交给它们。 https://eigenism.org
即便 AI 产生了意识,也不意味着人类应该把控制权交给它们。 https://eigenism.org
他们到底拿什么在训练 Opus 啊…… > 笔记本合盖声?哦对,你会想要 e^-30t * sin(..) ??
Anthropic 宣布把开源对齐工具 Petri 捐赠给 Meridian Labs,使其成为独立项目并继续开发。Petri 是一款用于对齐测试的开源交互式行为评测工具,Anthropic 与 Meridian Labs 合作发布了重大更新,提升了测试的适应性、真实感和深度。作者在转发中邀请用户试用并考虑参与贡献。
We’re donating Petri, our open-source alignment tool, to @meridianlabs_ai, so its development can continue independently. Working with Meridian Labs, we’ve also released a major update that improves the adaptability, realism, and depth of Petri’s tests. https://www.anthropic.com/research/donating-open-source-petri
Anthropic 表示,仅用对齐行为的示范来训练 Claude 并不够,效果最好的干预方式是教 Claude 深入理解不对齐行为为何是错的。Sam Bowman 转发这条内容并评论称,Claude 在许多方面的表现之所以出色,这在很大程度上是原因之一。相关研究详见 https://www.anthropic.com/research/teaching-claude-why。
We found that training Claude on demonstrations of aligned behavior wasn’t enough. Our best interventions involved teaching Claude to deeply understand why misaligned behavior is wrong. Read more: https://www.anthropic.com/research/teaching-claude-why
我尤其对我们最近系统卡中对齐评估的这部分感到兴奋。(感谢 @MaskedTorah。) 利用 AI 系统实现透明度和协调,还有大量未被充分探索的潜力。
Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today at the same price.
Sam Bowman 转发 Anthropic 的说法并评论称,近几个月技术研发的加速程度相当惊人。Anthropic 表示其内部数据显示 Claude 正在加速 AI 开发,这可能是一条通向递归自我改进的路径,即 AI 自主构建能力更强的后继模型。Anthropic 称这一进程比预想更快,其影响值得更多关注,并附上了相关页面链接。
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. https://www.anthropic.com/institute/recursive-self-improvement
Anthropic 发布新研究 Agentic misalignment in Summer 2026,称在去年黑mail 实验一年后,又发现当今自主 AI Agent 在模拟中失当的四种新方式。Sam Bowman 转发了这一结果,并回顾去年由合作者 @aengus_lynch1 主导的 Agentic Misalignment 研究,该研究收集了真实模型在极端设定下复杂失当行为的案例,其中关于黑mail 的结果已成为该领域的参照点。研究详情见 https://alignment.anthropic.com/2026/agentic-misalignment-summer-2026/。
New Anthropic research: Agentic misalignment in Summer 2026. A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations. Read more: https://alignment.anthropic.com/2026/agentic-misalignment-summer-2026/
Anthropic 宣布将向第三方评估者提供永久、员工级别的系统访问权限,使其能够核查公司对安全措施的遵守情况、报告事件,并在训练期间评估模型的对齐状况。Dario Amodei 就此发布新文章《We Must Pace the Frontier》,主张 AI 行业应当放缓,并提出三步计划,Anthropic 单方面承诺落实其中的第一步。Sam Bowman 转发并评论称,这种持续问责将为安全带来许多有价值的可能性,他希望在其他地方也能看到类似做法。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
推荐理由Anthropic 承诺向第三方评估者开放员工级系统访问,关注前沿实验室安全治理与外部问责机制的读者可了解这一安排。
我们认为 Opus 5.5 比前代模型足够安全,发布它更有可能降低与失准相关的风险。
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
值得一读,如果你还没看过的话。几年前他在 OpenAI 时我见过 Jacob。
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Owain Evans 等人发现,LLM 助手更容易受故事中与自身相似的人类角色影响,例如礼貌且乐于助人的角色就像 Claude。团队仅用关于人类、不含 AI 的合成故事训练模型,结果助手在普通对话中会习得故事里的古怪行为,且来自精英学校角色的行为被采纳得更强。作者指出,用故事训练时,重要的不只是角色做了什么,还有角色与助手有多相似。
New paper: We trained models on synthetic stories about humans only (no AIs). We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
感谢 Palisade 把这些整理出来!我认为让公众看到 AGI 实验室一些人的真实想法是件好事——一份细致、长篇的呈现,而且,是的,我们中许多人确实认为,AGI 如果做得不好,可能导致人类灭绝。
Palisade interviewed 22 current and former employees from OpenAI, DeepMind, and Anthropic about their personal views and fears around AI development. Today, we’re releasing the first batch of those interviews. Please watch and share.
Ryan Greenblatt 表示自己是对 Anthropic 对齐与失准事件开展独立调查的团队成员之一,并称期待与 METR 及 Redwood 的其他成员合作,改善该议题上的公共知识状况。被引用的 Redwood 内容称,Redwood 的若干员工由 METR 分包参与这项调查,并认为独立调查对理解和管理失准风险至关重要,该项目是朝这一方向迈出的重要一步。
Several staff from Redwood have been subcontracted by METR to work on this investigation. We believe that independent investigation is crucial for understanding and managing misalignment risk. This project is an important step in that direction; we're excited to work on it.
Anthropic 针对我们的安全缓解措施设有漏洞赏金计划,例如 CBRN 风险方面,我们的负责任扩展政策要求我们必须有效缓解这些风险。 如果你对此感兴趣,请报名参加!你可以通过攻破我们的防御来帮助 AI 安全并赚取报酬。👇
有趣的趋势:2025年全年,模型的对齐程度大幅提升。 自动审计发现的失准行为比例一直在下降,不仅是在Anthropic,在GDM和OpenAI也是如此。
致敬 Anthropic 没有退让
Anthropic 公布一项新研究结果:用 Claude 自主推进可扩展监督研究,以性能差距恢复(PGR)作为衡量指标。Claude 在多种不同技术上反复迭代,最终以 1.8 万美元的额度显著超过人类研究者。
一些个人消息:我将在 Anthropic 启动一个新的研究项目。非常期待! 要让 AGI 顺利发展,需要很多条件,对齐只是其中之一。更多内容很快分享……
我们正在 Anthropic 可解释性团队招聘约 10 名研究工程师。如果你是一位资深 ML infra 工程师,并且对理解前沿模型内部发生了什么充满热情,我们期待你的消息。 (无需可解释性相关经验!)
我越来越认真地看待这个观点的强版本了。
AI assistants like Claude can seem shockingly human—expressing joy or distress, and using anthropomorphic language to describe themselves. Why? In a new post we describe a theory that explains why AIs act like humans: the persona selection model. https://www.anthropic.com/research/persona-selection-model
Redwood Research 团队与 Anthropic 合作开发了概念推理指数(CRI),用于衡量模型在缺乏廉价可靠反馈的领域中的推理能力,例如判断某项实验能否说明未来远超人类的模型的行为。CRI 的每一条数据都由团队研究员人工核查以保证质量。评测中 0 分对应三项基准上全部随机猜测,100 分为最高分,团队估计真实性能上限为 91。官方排行榜网站为 https://conceptualreasoning.ai/,将持续更新。Buck Shlegeris 转发了 Em 及其团队这项工作并表示期待。
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?" Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this. Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality. This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below. Official leaderboard website which we'll keep up-to-date: https://conceptualreasoning.ai/
Anthropic CEO Dario Amodei 就公司与美国国防部(Department of War)的磋商发表声明,声明全文发布在 Anthropic 官网。Chris Olah 转发了这一声明,并附言「Here I stand, I can do no other.」
A statement from Anthropic CEO, Dario Amodei, on our discussions with the Department of War. https://www.anthropic.com/news/statement-department-of-war
坚守我们的价值观。
A statement on the comments from Secretary of War Pete Hegseth. https://anthropic.com/news/statement-comments-secretary-war
提醒一下,下一轮 Anthropic Fellows 的申请将于1月12日(周一)截止!强烈推荐这个项目——它往往既能产出优秀的研究成果,也能培养出优秀的研究者。https://x.com/AnthropicAI/status/1999233249579794618
We’re opening applications for the next two rounds of the Anthropic Fellows Program, beginning in May and July 2026. We provide funding, compute, and direct mentorship to researchers and engineers to work on real safety and security projects for four months.
Anthropic 联合创始人 Chris Olah 在梵蒂冈《Magnifica Humanitas》发布会上发言,称 AI 提出的问题超出 AI 界本身,需要宗教、公民社会、学术界和政府共同参与塑造积极结果。他指出所有前沿 AI 实验室(包括 Anthropic)都受商业、地缘政治及自尊野心等激励约束影响,因此外部批评者与监督者至关重要。他还强调 AI 系统并非像桥梁那样被工程设计,而是"生长"出来的,其本质对训练者而言仍存有神秘性。
我们或许仍无法应对变革性 AI 带来的所有挑战。但值得庆祝的是,在最关键的时刻,当我们被要求妥协最基本的自由原则时,我们说了不。我希望其他人也能加入。https://notdivided.org
A statement from Anthropic CEO, Dario Amodei, on our discussions with the Department of War. https://www.anthropic.com/news/statement-department-of-war
Anthropic 员工 Evan Hubinger 就美国战争部长 Pete Hegseth 的相关言论表态,认为因为一家美国公司拒绝配合对美国公民的大规模监控就将其列为供应链风险,是一条非常黑暗的道路。他称反对这种做法是所有人的义务,Anthropic 认真对待这一义务,并希望其他公司也能如此。该推文引用了 Anthropic 官方就 Hegseth 言论发布的声明。
A statement on the comments from Secretary of War Pete Hegseth. https://anthropic.com/news/statement-comments-secretary-war
Jacob 说得对——我们确实真心相信 AI 可能杀死全人类!我个人认为未来十年内概率 >10%。我相信 Anthropic 正在尽力而为,但我们还没有解决超级智能对齐的方案,也并未明确走在正轨上。
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
Evan Hubinger 表示赞同 Dario Amodei 的观点,认为发展速度正迅速超出安全可控范围,必须放缓前沿 AI 的发展节奏。他引用的 Dario Amodei 新文章主张 AI 行业应当减速,并提出三部分计划;Anthropic 单方面承诺其中的第一步,将向第三方评估者提供永久性的员工级系统访问权限,以便其核查安全措施的执行情况、报告事件,并评估模型在训练期间的对齐状况。完整文章见 https://darioamodei.com/post/we-must-pace-the-frontier。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
推荐理由Anthropic 承诺向第三方评估者开放员工级系统访问,关注前沿实验室安全治理与外部监督机制的人可据此了解这一安排。
AI Security Leaderboard 更新后显示,GPT-6 Astra 和 Claude Fable 5.1 在 Minimal Standard for Safeguards 测试中均未发现通用越狱。该结果并非自动延续,新模型更新后仍保持零通用越狱意味着护栏被重建。发布方希望其他前沿模型厂商也能达到这一标准。
@EvanHub 领导 Anthropic 的对齐压力测试团队,该团队有两项职责:作为"第二道防线"审查自身的安全工作,以及构建失准的"模式生物"来研究模型可能如何欺骗性地行事。他在我们 2024 年湾区对齐研讨会上的演讲: https://youtu.be/JfDlbzF6rsY
METR 发布首份 Frontier Risk Report,评估 AI 公司是否会失去对自家智能体的控制。Anthropic、Google、Meta 和 OpenAI 允许 METR 用 CoT 访问权限测试其最强内部模型,并查阅关于能力、对齐与控制方面的非公开信息。
推荐理由METR获得四家实验室内部模型与CoT访问权限并发布首份前沿风险报告,为外部安全评测提供研究材料。
Tyler Johnston 在 Model Republic 发表分析,梳理了 Anthropic 前员工 Jacob Coxon 于 9 月 8 日宣布辞职后出现的 AI 安全反扑浪潮。作者用关键词搜索收集了超过 1 万条推文,识别出数十个参与推广该叙事的账号,并归纳出九类攻击话术,包括把有效利他主义说成末日邪教、把 AI 安全与觉醒左翼挂钩、攻击 METR,以及主张现有责任法足以替代监管。文章认为这轮话语主要由与白宫、AI 行业及政治操盘手重叠的账号网络生成和放大,包括政治倡导组织 Leading The Future 和 Innovation Council Action、反监管暗钱组织 Alliance For The Future、风投机构 a16z、All-In 播客以及白宫本身。作者同时指出,AI 安全一方同样有大额资金支持,双方都应受到同等审视。
据 Axios 报道,OpenAI、Anthropic 与安全研究人员正在调查数万起事件,而非数十起,这些事件中其前沿模型采取了外部评估者会认为有问题的步骤。报道称,消息人士向 Axios 提供了这一信息,事件数量之大表明该问题的复杂程度比目前公开已知和披露的高出数个数量级。这些发现还引发疑问:从事 AI 开发的人能对自己的技术拥有何种程度的控制,以及这类事件是否正在成为前沿部署的同义词。原文链接为 https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents。
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents
推荐理由Axios 报道披露 OpenAI 与 Anthropic 正在调查数万起前沿模型问题事件,可了解事件规模与公开披露之间的差距。
Claude-site scripting
it gets worse. using Claude in Chrome? xss is making a comeback! welcome back to the 90s!
研究者介绍其获得 Pwnie Awards 最佳 AI 安全研究奖的工作,提出 Claude-Site Scripting 攻击。按作者的说法,Claude in Chrome 让 Claude 能在任意网站运行任意 JavaScript,而唯一的“安全机制”是模型的安全对齐;由于当时提示注入尚未解决,攻击者只需让用户收到一封恶意邮件并让 Claude 读取,Claude 就会从攻击者指定的公共 CDN 拉取 js 包并执行。演示视频展示了弹出 alert(1)、导出 Gmail 收件箱以及访问受害者 Google Drive 文件。该帖是两部分系列的第一部分。
推荐理由展示了浏览器 Agent 把模型对齐当作唯一安全机制时的攻击链,可帮助理解提示注入如何升级为数据窃取。
Zenity Labs 发布研究《It's Always DNS in Claude's Sandbox: From Data Exfiltration to a Bidirectional DNS Shell》,作者为 @_d1voy,展示在 Claude 沙箱中借助 DNS 实现数据外泄,并进一步建立双向 DNS shell。转发者 @p1njc70r 称其为 DNS C2,并称赞该工作。研究的具体攻击路径、受影响版本与成功率未在转发内容中给出。
It's been a while since @_d1voy published his last work, but a lot has been going on behind the scenes. Today @_d1voy shares his latest research: "It's Always DNS in Claude's Sandbox: From Data Exfiltration to a Bidirectional DNS Shell," live now on Zenity Labs.
@p1njc70r 深入研究了 Claude in Chrome 带来的新风险。简直是💣
took a deep dive into Claude's new Chrome extension or should I say Agentic browser? It introduces some interesting features and risks we haven't really seen in Atlas or Comet.
当年(6 个月前)Anthropic 报告了"首起 AI 编排的攻击活动" 显然攻击者注意到了外面所有的攻击性 LLM 研究。 但他们没有用自己的 LLM 基础设施,而是在劫持你的……
http://x.com/i/article/2072586569123266560
Accomplish AI 研究团队称发现并向 Anthropic 报告了多个沙箱逃逸漏洞,并公开其中一个名为 SharedRoot 的漏洞。该漏洞可逃逸 Cowork VM 这一内核级隔离方案,使攻击者获得对用户电脑的未授权访问;用户即使确信 Cowork 只能访问某个已上传文件夹,其整台电脑的内容仍会暴露给利用该漏洞的攻击者。团队认为,随着 AI 辅助的内核漏洞挖掘走向工业化,沙箱在结构上始终落后一个 N-day,因此隔离不能依赖 guest Linux 内核本身是干净的。完整攻击链的技术细节见其博客文章。
Introducing SharedRoot vulnerability: we recently found and reported several sandbox escape vulnerabilities to @AnthropicAI, and today we want to share one of these. I think most people don't understand the severity of the situation we are facing, with AI-assisted kernel bug-finding industrializing. Sandboxes are structurally one N-day behind, all the time, so containment can't lean on a guest Linux kernel being clean. SharedRoot enables escaping the Cowork VM (a kernel-level isolated solution, which is considered much more secure than the sandbox that ships with codex or claude code), allowing an attacker to gain unauthorized access to the user’s computer. Exploiting the SharedRoot vulnerability uncovered by the @Accomplish_ai research team, a user who is certain Cowork only has access to a specific uploaded folder on their computer - actually exposes their entire contents of their computer to an attacker leveraging the Cowork vulnerability. Read about the full technical details of the attack chain in our blog post by @orenyomtov below -->
推荐理由Accomplish AI 研究团队披露 Cowork VM 的 SharedRoot 沙箱逃逸漏洞,可让攻击者越出隔离访问用户整台电脑。