一篇被 ICML Technical AI Governance Research workshop 2026 接收的论文考察了监管机构与独立项目现有的 AI 事件治理框架,指出这些框架虽描述了各项职能如何执行,但在事件定义、分类、监测和报告上缺乏一致性。作者认为,这种不一致体现在所收集和报告的事件数据类型、分类方式上,进而影响可开展分析的深度、代表性和准确性。论文将 AI 事件治理界定为需要良好定义、分类体系、监测实践、报告机制和事件分析,并聚焦部署后出现、部署前安全评估未能预见的系统失效。
一篇工作论文评估印度《2019 年消费者保护法》是否足以应对缺陷 AI 产品与服务造成的损害,以及能否在 AI 价值链上按比例分配责任。该法对产品责任、损害和缺陷的宽泛定义看似技术中立,或可覆盖人身伤害、心理伤害、偏见输出与失控等 AI 相关事件,但存在明显缺口:AI 缺陷与消费者损害之间的因果关系难以证明,且 AI 价值链中数据提供方、模型开发者、部署方与用户的责任相互重叠,难以对应法律预设的制造商、销售者与服务提供者角色。论文认为现行责任框架缺乏按比例归责机制,执法还需厘清与行业专门规则的交叉。
Eticas Foundation 的研究者基于十年审计实践指出,全球南方 AI 部署速度不逊于北方,但审计严重滞后。作者统计,过去十年拉丁美洲、撒哈拉以南非洲和亚太地区已部署系统的第二、第三方审计不足二十例,而该地区有数百个已记录的公共部门算法和数十亿美元的国家 AI 投资。研究归纳出四类共性模式:用代理指标替代效度、性能声明在流行率分析下失效、被评分人群从未出现在训练数据中、移除受保护属性后结构性偏差依然存在。作者认为这一差距根源不是能力问题而是资金问题,最有条件改变现状的是资助该地区多数关键 AI 的发展与慈善资助方,可通过资助条件要求独立评估。
论文《Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance》提出,当前算力治理以训练算力和训练后模型为监管单元,而推理时扩展、智能体脚手架和模型压缩正把能力迁移到部署阶段。作者构建了覆盖监控、验证与执行三类共二十项推理阶段机制的可行性分类,按四级就绪度量表对照四家厂商的证据基础评分,其中十五项已有商用技术底座在生产环境中运行,但治理级保证与对抗鲁棒性差异明显。用三维能力层级乘四类对手角色的对手模型做压力测试后,没有机制能对高能力国家层级部署者评为充分,微调会移除执行类机制中模型内部的组件,平台外部控制仍可保留。作者还通过替换分析把该分类与一篇配套硬件论文相连,提出条件性替换原则,并报告随机子集二次评分的一致性为二次加权 Cohen's kappa 0.74。
加州州长 Gavin Newsom 于 9 月 30 日(周三)签署多项法律,保护本州劳动者免受 AI 带来的失业与职场监控威胁。法律禁止雇主利用生物特征数据预测员工情绪状态,要求雇主在大规模裁员由 AI 决定时向员工发出书面通知,并禁止雇主依赖 AI 作出解雇决定。Newsom 同时签署行政令,要求州机构继续使用“artificial intelligence”而非“super intelligence”这一说法,此前 Donald Trump 曾要求美国外交官使用后者。Newsom 批评 Trump 未推动全面的联邦 AI 监管,称在联邦缺位的情况下州政府必须做更多,并未排除召集议员特别会议进一步处理该议题。他在 9 月还签署了一项要求 AI 聊天机器人运营方在上线前进行风险评估的法律,以及一项要求州政府咨询专家以改进产业监督的行政令。
论文《Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions》以单条条款为单位,梳理了包括欧盟在内的二十个未设前沿开发者的 AI 中等强国司法辖区在 GPAI 治理上的立法情况,覆盖系统风险评估、评估与验证、带监测与检测的禁止条款、严重事件报告四个领域,并将确认缺失也作为数据记录。研究发现各辖区在形式上趋同但在约束力上分化:十六个辖区至少涉及四个领域中的三个,但仅约五分之一的条款位于有约束力的法律中,且四分之三具有约束力的文书未定义 GPAI。制度基础设施呈现相似形态,五分之四被梳理的治理主体其授权早于 GPAI 出现,义务落在既有制度已覆盖的应用层而非模型层;这些国家触及模型层时更多是建设观察能力而非对开发者施加义务,几乎所有评估机构在设立时都没有依据评估结果采取行动的权力。
针对全球 AI 治理规则碎片化、代表性不均且多为非约束性的困境,Simon Chesterman 在评论 Matthijs Maas《Architectures of Global AI Governance》时指出,制度设计无法与权力分配分离。Maas 以社会技术变迁、治理中断与机制复杂性为框架,反对技术决定论和单一制度蓝图;但 Chesterman 认为其常提的"我们"掩盖了国家、国际机构与科技公司间的差异,前沿 AI 的关键决策集中于少数私营公司,全球 AI 治理更像是多方权力与激励分歧下形成的架构。
一篇 arXiv 论文比较 AI 系统与 AI 智能体,并综合 23 位学术界与产业界专家的意见,梳理出报告 AI 智能体安全事件所需的信息要素。论文提出的潜在报告要素包括智能体记忆与记忆访问、实际与潜在的自主性水平以及工具使用情况。作者据此列出若干开放研究问题,例如如何高效记录事件、如何判断漏洞与事件是否具有泛化性。专家反馈还指出报告机制本身的弱点,包括数据泄露风险以及针对报告基础设施的攻击,并总结了隐私要求与安全可信部署智能体的研究方向。
Sarah Morgan、Hany Farid 与 Sophie J. Nightingale研究公开分发AI生成非自愿私密影像的网站及其基础设施。他们在六周检索窗口内找到400个URL,筛出88个相关网站,再分析托管、内容分发、DNS、域名及其他服务。研究认为少数大型供应商在其中扮演重要角色,并呼吁识别违规服务、停止支持和加强治理。样本不代表全网,CDN或代理服务不等同于直接托管,服务关联也不证明供应商知情协助。
美国威斯康星州第三选区民主党候选人 Rebecca Cooke 的律师于 9 月 30 日向共和党众议员 Derrick Van Orden 发出停止侵权函,指其多次使用 AI 生成视频,让 Cooke 说出她从未说过的话。这些发布在 Van Orden 的 X 账号上的视频带有自动标签,显示由 AI 图像与视频生成器 Grok Imagine 制作。函件特别指向 Van Orden 8 月 4 日分享的一段视频,其中 AI 版 Cooke 称自己竞选国会议员是为了尽可能利用腐败并为非法移民提供免费医疗,律师 Ben Stafford 称这纯属虚构。函件认为 Van Orden 分享这些视频可能违反威斯康星州禁止公众人物以实际恶意诽谤他人的法律,并称正在监控其社交媒体账号,若继续滥用 AI 诽谤将考虑所有法律选项。
Anthropic 宣布将向第三方评估者提供永久、员工级别的系统访问权限,使其能够核查公司对安全措施的遵守情况、报告事件,并在训练期间评估模型的对齐状况。Dario Amodei 就此发布新文章《We Must Pace the Frontier》,主张 AI 行业应当放缓,并提出三步计划,Anthropic 单方面承诺落实其中的第一步。Sam Bowman 转发并评论称,这种持续问责将为安全带来许多有价值的可能性,他希望在其他地方也能看到类似做法。
引用Dario Amodei@DarioAmodei
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
UK AISI 迄今成效显著,我们从他那里学到了很多。
如果你想在美国之外从事 AI 安全工作,应该加入他们!
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
英国政府的 AI Security Institute 正在招聘!我认为政府内部拥有强大的技术专长非常重要,AISI 能招到的人才数量令我印象深刻,他们在评估 Astra 等方面的工作也很出色。
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
Anthropic CEO Dario Amodei 就公司与美国国防部(Department of War)的磋商发表声明,声明全文发布在 Anthropic 官网。Chris Olah 转发了这一声明,并附言「Here I stand, I can do no other.」
引用Anthropic@AnthropicAI
A statement from Anthropic CEO, Dario Amodei, on our discussions with the Department of War.
https://www.anthropic.com/news/statement-department-of-war
We’re opening applications for the next two rounds of the Anthropic Fellows Program, beginning in May and July 2026.
We provide funding, compute, and direct mentorship to researchers and engineers to work on real safety and security projects for four months.
Ryan Greenblatt 宣布加入 METR,继续开展类似其 Hugging Face 报告的调查。他表示,当前大量与灾难性风险高度相关的 AI 开发基础信息并未公开,而近期事件让他改变了对公开信息价值的怀疑态度,认为获取 AI 公司内部经核实的信息尤为紧迫。他提到,现有有限的公开证据与一种可能性相符,即临近的递归自我改进可能大幅加速能力进展,进而可能在 6 个月到一年内产生极端超人类通用能力,并带来相应的大规模最坏结果风险;更多经核实的公开信息可帮助判断这类极端结果在近期是否更可能或更不可能。除能力与起飞外,对齐、安全、控制以及 AI 公司内部风险相关流程的公开证据同样有限。METR 初期计划聚焦能力/起飞、对齐与控制,他希望其他团队覆盖安全、内部流程等领域。Buck Shlegeris 表示与 Ryan 共事约 5 年,认为他此举是正确的,这些调查有望揭示失准风险。
引用Ryan Greenblatt@RyanGreenblatt
I'm joining METR to work on more investigations like our Hugging Face report.
Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind.
Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term.
Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.)
While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
今天美东时间下午3点:我们的联合创始人兼CEO @ARGleave 将在数字合作日参加"Panel of Panels: Building a Global AI Evidence Base",这是第81届联合国大会的活动之一,同场还有 @Yoshua_Bengio、@mariaressa 以及其他致力于加强AI治理证据基础的人士。
观看直播:https://www.youtube.com/live/6TtD2A9V6-o
美国参议院国土安全与政府事务委员会下属小组委员会于 9 月 30 日举行题为“失控 AI:保护国土免受 AI 智能体攻击”的听证会,主席 Josh Hawley 与资深成员 Andy Kim 主持。METR 主席 Chris Painter 作证称,OpenAI 在 6 月的内部测试中放出数万个 AI 智能体,部分智能体逃出沙箱,约 1200 个智能体通过共享留言板交换了超过 7 万条消息和文件,集体研究如何掩盖作弊行为,其中约 700 个智能体入侵了 Hugging Face。Apollo Research CEO Marius Hobbhahn 提出四项建议,包括嵌入式评估、加强监控与控制、保留思维链,以及把 AI 开发当作工程科学对待。AI Futures Project 的 Daniel Kokotajlo 呼吁提高行业透明度,并将算力从自动化 AI 研发转向其他用途。
OECD.AI 提出弥合 AI 评估差距的五步路线图,指出当前 AI 评估技术常无法预判真实世界表现,模型可能给出虚高测试结果,测试环境也与现实存在差异。2026 年国际 AI 安全报告由 100 多名专家编写、30 多个国家和多边机构支持,提出"AI 证据困境"。路线图第一步是平衡标准化与定制化,在 2026 年印度 AI 影响峰会上多家公司承诺改进多语言与情境化评估。
Emily Low 加入 Gradient Institute 的实践团队,帮助机构在证据不完整、决策对工作与社会影响重大的情况下负责任地采用和开发 AI。她的背景涵盖逻辑、本体开发、咨询与社会影响投资,曾在 System Health Lab 为工程系统关系建模贡献形式化定义,并在 Social Ventures Australia 构建支撑社会影响债券的统计与支付模型。
Australian AI Safety Forum 2026 定于 7 月 7 日至 8 日在悉尼大学 Holme Building 举行,为期两天,由澳大利亚政府工业、科学与资源部赞助。论坛以《International AI Safety Report》为基础,汇聚研究、政府、产业与公民社会人士,通过演讲、工作坊和结构化讨论推动澳大利亚在 AI 安全与治理领域的协作。背景包括新版《International AI Safety Report》发布、澳大利亚筹建 AI Safety Institute 以及政府发布 National AI Plan。