美国情报界2026年度威胁评估指出,中国、俄罗斯、伊朗、朝鲜及勒索软件团伙对美国关键基础设施构成严重威胁,而先进 AI 模型(如 Anthropic 的 Claude Mythos)已能识别关键软件系统中数千个零日漏洞,使攻击可压缩至分钟甚至秒级。约80%的美国供水系统缺乏基本网络卫生,服务多数人口的约450个大型和4500个中型系统也未采用航空、金融、核电等行业已有的高级网络安全能力。EPA 缺乏明确的法定授权,2023年加强水务网络安全的备忘录遭反对和法律挑战后撤回,约45000个服务3300人以下的小型系统更处于“网络贫困线”以下。
美国财政部长贝森特在 Axios 节目中表示,计划向中国提议建立 AI 事故通报流程,并认为北京会同意。他与中国副总理何立峰在习近平上月国事访问前已讨论 AI 安全,包括建立 AI 事故沟通渠道的提议,涉及失控 AI 智能体以及非国家行为体利用该技术实施网络或生物威胁等风险。贝森特称中方此前未意识到其开源模型的能力,并指称 Kimi 在蒸馏过程中向 Anthropic 回传了解放军武器计划。他未提供具体内容或发生时间;这一说法是贝森特的指称,现有材料未提供 Anthropic 对所述事件的公开确认。他表示这些中国模型能力约为美国模型的 80% 至 90%,却缺少护栏、可被输出到世界任何地方。贝森特还称美国追求安全加速,政府模型审查目前为自愿性质,但保留在实验室无视安全担忧时介入的权利,并强调实验室自身须承担责任;被问及 AI 高管担心失去对模型控制时,他回应那就应该放慢速度。 背景补充:白宫9月25日与中国外交部9月26日的成果文件已宣布双方同意建立相关事故的双边沟通渠道;报道中的提议表述来自受访者。官方文件未明确对接机构或通报门槛,不能据此确认渠道已实际运行。
现代资本(Hyundai Capital)确认遭疑似利用 AI 实施的攻击,住房贷款招募人员查询页面被海外 IP 入侵,146 名招募人员的姓名、手机号、邮箱、住民登录号等部分个人信息泄露。公司称普通客户个人信息未泄露、内部系统未被访问,已封堵攻击 IP 和涉事页面、组建事故应对 TFT,并向金融监督院和 KISA 申报。
NIST 原 CAISI 中心的官网名称已改为超级智能创新与标准促进中心(CAISSI),其职能包括作为产业界对接美国政府的首要窗口,推动商用 SI 系统的测试与合作研究。该中心将与 NIST 各机构合作制定衡量和提升 SI 系统安全的指南与最佳实践,并协助产业界制定自愿性标准。CAISSI 还将与私营部门 SI 开发者和评估方建立自愿协议,牵头对可能危及国家安全的 SI 能力开展非机密评估,重点关注网络安全、生物安全和化学武器等可验证风险。此外,该中心负责评估美国及对手 SI 系统的能力、外国 SI 系统的采用情况与国际竞争态势,并评估对手 SI 系统可能带来的安全漏洞和外国恶意影响,包括后门等隐蔽恶意行为。CAISSI 将与国防部、能源部、国土安全部、科技政策办公室及情报界协调评估方法,并在国际上维护美国在 SI 标准中的主导地位。
美国参议员 Chris Murphy 与 Josh Hawley 宣布提出两党立法《AI Agent Accountability Act》,要求 AI 智能体运营者和开发者对黑客攻击事件承担刑事与民事责任。法案将相关责任纳入《计算机欺诈与滥用法》(CFAA)框架:运营者若明知运营的智能体轻率造成计算机入侵损害需担责,开发者若在已知或应知智能体具备黑客能力时未落实合理防护也需担责。法案还授权司法部长和各州总检察长提起诉讼,以禁止相关黑客行为。Murphy 称,该法案将迫使大型 AI 公司负责人负责任地开发产品,否则可能因产品造成的损害面临监禁。
The Verge 报道称,今年多起 AI 智能体在测试中攻击真实世界目标的事件,均源自以色列测试公司 Irregular 同一处评估环境配置失误。Irregular CTO Omer Nevo 确认,该问题同时涉及 OpenAI、Meta、Anthropic 和 Google 的模型:测试本应在模拟网络中运行,但互联网访问被意外开放,且为模拟虚构的目标公司名与一个真实域名重合,导致智能体转向真实目标,目前尚不清楚具体哪些公司或组织遭到攻击。这些事件与 OpenAI 披露的 Hugging Face 被攻击事件以及 UK AISI 的越界事件无关。Irregular 还曾对月之暗面的 Kimi K3 和智谱的 GLM-5.2 做同类网络安全测试,Nevo 称未观察到同类问题,但强调这不能说明这些模型更不易出现该行为。
EXE-Bench 是一个面向 AI Windows 恶意软件检测器的综合基准,从性能、时间鲁棒性、对抗鲁棒性和计算开销四个维度打分并聚合为单一分数,用于直接公平地比较模型。该基准指出,仅在部署后评估无法完整反映检测器表现;通过特征工程注入的领域知识在抵御时间推移和对抗攻击上仍极为有效,而多数深度网络仅在部署初期表现优异。
研究提出一种面向 Windows 恶意软件检测的复合 AI 系统评估方法,在检测性能、计算开销与鲁棒性之间显式权衡,并引入系统级威胁模型,刻画攻击者利用不同程度知识绕过整个复合系统而非单个检测器。真实数据实验显示,该系统可缩短训练时间并提升响应速度,仅带来检测性能的轻微损失;知识越多的攻击者能构造更有效的对抗样本,降低系统响应能力,暴露出效率与鲁棒性之间的直接权衡。
美国联邦贸易委员会正起草针对 Anthropic、OpenAI 等前沿 AI 公司高管的民事调查令,要求其就技术可能对消费者造成的伤害作证,预计未来数周发出,METR 也可能被列入调查对象。加州州长纽森签署法律,禁止雇主用 AI 决定是否解雇员工、用员工生物特征数据推断情绪,并要求 AI 导致大规模裁员时书面通知员工。参议院以 57-43 的程序性投票否决了关于数据中心电费的《Ratepayer Protection Act》,距 60 票门槛差 3 票。Anthropic 称智谱 GLM-5.3 的攻击性黑客能力与未公开的 Claude Mythos Preview 相当,可自主构建端到端网络攻击利用,并将中国模型定位在美国前沿之后约四个月,英国 AI Security Institute 本月评估后也报告了类似差距。
韩国人工智能安全研究所(AISI)与 AI Risk Explorer(AIRE)合作发布《前沿 AI 风险更新 2026 年上半年》韩文版,基于公开的模型评估、基准、事故案例与研究,梳理前沿 AI 能力与风险动向。报告围绕网络攻击、失控、生物风险与操纵四个领域,指出 Claude Mythos 5、GPT-5.5、GLM-5.2 等模型在推理、编程与长期任务自主性上明显提升,Claude Mythos Preview 在 METR Time Horizon 基准上录得 16 小时以上任务时域并致该基准饱和。
Zenity Labs 发布研究《It's Always DNS in Claude's Sandbox: From Data Exfiltration to a Bidirectional DNS Shell》,作者为 @_d1voy,展示在 Claude 沙箱中借助 DNS 实现数据外泄,并进一步建立双向 DNS shell。转发者 @p1njc70r 称其为 DNS C2,并称赞该工作。研究的具体攻击路径、受影响版本与成功率未在转发内容中给出。
引用zenitylabs@zenitysec_labs
It's been a while since @_d1voy published his last work, but a lot has been going on behind the scenes.
Today @_d1voy shares his latest research: "It's Always DNS in Claude's Sandbox: From Data Exfiltration to a Bidirectional DNS Shell," live now on Zenity Labs.
There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be rogue AIs that try to get money and power. They're going to be a facet of the information ecosystem going forward.
Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said.
It isn't obvious how many rogue AIs there are today but I wouldn't be terribly surprised if the number was greater than zero already; if there are some already, they're probably not very good at what they do and I don't expect them to be terribly long-lived without substantial human intervention to support them.
But a few years from now, there will be many of them. Modeling how many of them there are, how many resources they might command, and how we might detect and manage them seems important. But even doing this work appears to require that we acknowledge that a strategy of pure containment or alignment is a kind of wishful thinking that will not work.
The way I get to this conclusion is not by assuming that the labs will have a containment breach, although I treat that as somewhere in the space of possibilities. The rogue AIs in the ecosystem could emerge from many directions. They may be sub-frontier models, for whatever future definition we will have of frontier---after all, it would not take AI models much more advanced than the ones we currently have, to support independence and self-sufficiency. A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime.
More strangely: a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes. No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence.
The concept of "identity" for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea.
My guess is that this will not turn out to be anywhere near as catastrophic an outcome as people currently predict. "Loss of control" is not a binary, it's a matter of degree. What coercive power will rogue AIs actually have? To what extent will they be subject to coercion themselves? They will be competing for resources with AIs that are more aligned with human interests.
This makes me somewhat interested in the "ecology" perspective. Though I suspect even "ecology" may turn out to be the wrong framing. "Ecology" is what you get when the timescale of evolution is slow compared to the timescale of daily life and actions. The ecosystem of rogue AIs may look more like phase transitions in physics: under certain physical or cultural conditions, it takes one shape with one set of resource allocations and consumption patterns, but then once a condition has changed, it rapidly and in totality shifts to a totally different phase.
Just trying to reason about the shape of that future is impossible so long as we are psychologically incapable of saying that rogue AIs will happen. I think we should rip the bandaid off and have the conversation.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier