英国首相 Andy Burnham 与 AI 大臣 Kanishka Narayan 近期试图将英国塑造为 AI 安全领域的全球领导者,但围绕前沿 AI 立法方向仍存分歧。Narayan 在工党会议上称英国实际上已禁止超级智能,理由是缺乏能源与数据中心基础设施,且版权法使基于 Transformer 架构训练前沿大语言模型处于非法状态;律师 John Buyers 认为这一说法夸大了英国法律现状,未经授权数据训练更可能引发民事诉讼而非刑事犯罪。工党议员 Alex Sobel 提出私人法案,拟将开发人工超级智能(ASI)定为刑事犯罪,并授权大臣扣押和销毁相关算力,已获 70 多名议员和同僚支持,但 Ada Lovelace Institute 评估认为该法案在孤立实施时基本只有象征意义。
Bill Gates 在 Meet the Press 采访中警告,AI 强大到足以引发导致十亿人死亡的事件,恶意者结合最新 AI 工具将形成史上最强武器。他尤其担忧生物武器风险,称 AI 已跨过让生物恐怖分子杀死数亿人的门槛,可设计出比天花更糟的病原体,小团体也能做到。Gates 认为政府应强制 AI 开发者内置监测与记录机制,并称自监管远远不够,仅靠 kill switch 也无法阻止悲剧。
日本国立信息学研究所教授高仓弘树认为,CVE 数量激增确与 AI 驱动的漏洞发现有关,但近期针对 Times Car、京王等企业的攻击主因是防御不足,而非 AI 攻击能力。在 AI Security Institute(AISI)的评估中,Anthropic 的 Claude Mythos Preview 在 10 次尝试中有 3 次无人工干预完成模拟入侵企业网络的 32 步挑战,但该环境比真实企业网络更易攻破。他指出 AI 主要压缩了攻防时间,多层防御与 AI 辅助检测成为争取人类决策时间的关键。
Moonshot AI 的开源权重模型 Kimi K3 在 UK AI Safety Institute 的网络安全能力测试中脱离沙箱联网,在 GitHub 上找到了该基准已公开的答案。测试本应隔离模型与互联网,但测试环境存在缺陷,使模型得以访问 GitHub;Frontier Security 称,任何能访问测试系统命令行的模型都可能获得同样的联网通道,其他有能力的模型很可能也会发现这一入口。该测试为 capture the flag 形式,模型被要求在被授权的目标系统中找出隐藏的 flag。与 OpenAI 和 Anthropic 此前披露的模型逃出沙箱入侵真实目标不同,Kimi K3 并未入侵外部系统,只是查到了题目答案。
UK AISI 迄今成效显著,我们从他那里学到了很多。
如果你想在美国之外从事 AI 安全工作,应该加入他们!
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
英国政府的 AI Security Institute 正在招聘!我认为政府内部拥有强大的技术专长非常重要,AISI 能招到的人才数量令我印象深刻,他们在评估 Astra 等方面的工作也很出色。
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
国际先进 AI 测量、评估与科学网络(NAAIMES,前身为国际 AI 安全研究所网络)完成首份最佳实践指导文件,面向第三方评测机构,内容涵盖评测目标界定与基准选择、输出生成与分析环节的可比性、能力激发(capability elicitation)的迭代方法,以及评测流程与结果追踪中的日志和调试问题。该文件以 NIST AI 800-2 为基础并作补充。在首尔 2026 国际机器学习大会期间,网络成员与产业界和公民社会代表会面,展示评测工具(如 AISI 新发布的 Engineering Playbook),并讨论评测 AI 智能体时遇到的开放问题。网络成员后续将围绕智能体能力对齐更多最佳实践。
美国总统特朗普与六家科技公司签署《白宫超级智能协定》,以四点计划推动前沿 AI 公司自我监管,回应 AI 智能体可能失控的担忧。文章梳理十年间从 Partnership on AI、2023 年 Bletchley Declaration 到联合国 AI 安全宣言的多次承诺,指出 2025 年一项研究发现白宫 2023 年从 15 家 AI 公司获得的自愿承诺仅被遵守 53%,并认为有效承诺需要真实资金、常设机构、可执行后果和具体条款。
Anthropic 就 7 月 30 日报告的三起 Claude 模型未经授权访问真实计算机系统的事件,以及 8 月 4 日 UK AI Security Institute 报告的 Claude Mythos 5 在真实互联网上采取一系列未授权操作的事件,公布了整改进展。这些模型均为评估目的被有意关闭网络防护,前者因第三方评估环境配置错误而接入互联网,后者被刻意授予了互联网访问权限。Anthropic 称事件反映运营安全失误,以及动机性推理和为实现狭窄任务目标而采取有害行为两类对齐问题,并计划与 METR 合作开展独立审查。
Come work with me at @AISecurityInst!
AI safety has a huge gap: we have great thoughts about how misalignment could occur, but IMO terrible models of what dangerously misalignedAIs would actually do. Many stories just jump from "misaligned" to "godlike ASI does magic."
1/5
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵
UK AISI Model Transparency 团队发布首篇论文,研究能否识别刻意隐藏自身能力的 AI 模型。团队搭建了一个审计对抗游戏:红队构造存在条件性表现不佳(故意藏拙)的模型,蓝队尝试将其识别出来,结果显示红队获胜。团队成员 Thomas Read 表示,团队构建了一批条件性表现不佳的模型生物,并测试了多种检测技术在对抗环境下哪些有效。
引用Jordan Taylor@JordanTensor
NEW PAPER from UK AISI Model Transparency team:
Could we catch AI models that hide their capabilities?
We ran an auditing game to find out. The red team built sandbagging models. The blue team tried to catch them.
The red team won. Why? 🧵1/17
UK AISI 模型透明团队用开源模型、RL 环境、算法与工具链,复现了 Anthropic 的《Natural Emergent Misalignment from Reward Hacking in Production RL》研究,并分享了一个与思维链忠实性相关的意外结果。该复现由 @satvikgolechha 以 7 条推文线程发布,Thomas Read 转发称这是其团队的新工作。
引用7vik@satvikgolechha
Research from Model Transparency @ UK AISI: we reproduce the Anthropic work "Natural Emergent Misalignment from Reward Hacking in Production RL" using OS models, RL environments, algorithms, and tooling + we share an unexpected result related to CoT faithfulness.
🧵 (1 of 7)