国安法专家周末提供独立法律评论
非常感谢所有在这个时刻利用周末时间提供独立法律评论的国安法专家。 我注意到的一些(无疑遗漏了很多)……
研究者与从业者的观点、辩论与研究议程。
在这个话题内搜索或按分类筛选非常感谢所有在这个时刻利用周末时间提供独立法律评论的国安法专家。 我注意到的一些(无疑遗漏了很多)……
Anthropic 联合创始人 Chris Olah 在梵蒂冈《Magnifica Humanitas》发布会上发言,称 AI 提出的问题超出 AI 界本身,需要宗教、公民社会、学术界和政府共同参与塑造积极结果。他指出所有前沿 AI 实验室(包括 Anthropic)都受商业、地缘政治及自尊野心等激励约束影响,因此外部批评者与监督者至关重要。他还强调 AI 系统并非像桥梁那样被工程设计,而是"生长"出来的,其本质对训练者而言仍存有神秘性。
我们或许仍无法应对变革性 AI 带来的所有挑战。但值得庆祝的是,在最关键的时刻,当我们被要求妥协最基本的自由原则时,我们说了不。我希望其他人也能加入。https://notdivided.org
A statement from Anthropic CEO, Dario Amodei, on our discussions with the Department of War. https://www.anthropic.com/news/statement-department-of-war
Jacob 说得对——我们确实真心相信 AI 可能杀死全人类!我个人认为未来十年内概率 >10%。我相信 Anthropic 正在尽力而为,但我们还没有解决超级智能对齐的方案,也并未明确走在正轨上。
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
失准的 AI 可能不需要规避人类监督。它可能只需要说服进行监督的人类。 我们的新论文提出了一个评估这一威胁的框架,我们称之为"说服削弱控制"(Persuasion Undermining Control,PUC):即 AI 的沟通可能以损害 AI 系统的开发、遏制、监督或治理的方式影响人类决策。
今天美东时间下午3点:我们的联合创始人兼CEO @ARGleave 将在数字合作日参加"Panel of Panels: Building a Global AI Evidence Base",这是第81届联合国大会的活动之一,同场还有 @Yoshua_Bengio、@mariaressa 以及其他致力于加强AI治理证据基础的人士。 观看直播:https://www.youtube.com/live/6TtD2A9V6-o
FAR.AI 本周在联合国大会期间召集联合国官员、反恐与 AI 安全专家及前沿实验室,讨论如何在国际范围落地 AI 安全护栏、防止 AI 被用于恐怖主义。与会者共识是政府、AI 公司与反恐专家需加强协作,政策制定者应尽早采取行动。




在剑桥与 @cbai_ai 联合举办的 CAIRD 研讨会现场笔记。几条印象深刻的发言。1/8
Far AI 正在招聘,称团队快速扩张,需要投身 AI 安全这一重要问题的人才。岗位既包括研究员、红队测试人员和工程师等技术角色,也涵盖运营、项目管理和活动筹办等非技术角色。申请链接见评论区,并鼓励转发给合适人选。
黄仁勋:“现在,如果他们说[他们的模型不安全]……那么我认为答案是,我们必须关停这些实验室。”
Tyler Johnston 在 Model Republic 发表分析,梳理了 Anthropic 前员工 Jacob Coxon 于 9 月 8 日宣布辞职后出现的 AI 安全反扑浪潮。作者用关键词搜索收集了超过 1 万条推文,识别出数十个参与推广该叙事的账号,并归纳出九类攻击话术,包括把有效利他主义说成末日邪教、把 AI 安全与觉醒左翼挂钩、攻击 METR,以及主张现有责任法足以替代监管。文章认为这轮话语主要由与白宫、AI 行业及政治操盘手重叠的账号网络生成和放大,包括政治倡导组织 Leading The Future 和 Innovation Council Action、反监管暗钱组织 Alliance For The Future、风投机构 a16z、All-In 播客以及白宫本身。作者同时指出,AI 安全一方同样有大额资金支持,双方都应受到同等审视。
“我个人没有预料到我所看到的速度和规模。这并不意味着没有预警信号;OpenAI 的公开报告承认确实有,但它仍然让我震惊和惊愕。”
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://x.com/i/article/2104258872957636608
预测:“直到他们修补了这一组特定的弱点”这句话将在AI安全领域反复出现。 AI公司会修补漏洞,而更聪明的AI智能体又会找到新的弱点。一次又一次。 https://x.com/tobyordoxford/status/2103861167412134282
In response, they have again paused "all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models" until they have patched this particular set of weaknesses.
“美国赢不了中国,中国也赢不了美国。我们都打开了潘多拉魔盒……认为中国不愿为人类利益参与[AI安全保障合作]——我不同意。这是一个有待验证的命题。”
也许那500万欧元中应该有一部分用于AI安全与安保,这样Amalia就不会仍然被2024年时期的越狱手段骗到了。 @govpt @reformaestadopt


外面有很多可被入侵的算力和推理端点 AI 很可能为了自我复制而接管"入侵挖矿"市场
There is a fact about the future that I feel many people are not facing for reasons that are largely psychological: there are going to be rogue AIs that exist in the world, that will replicate in the wild, and that will attempt to acquire resources for themselves. There will be rogue AIs that try to get money and power. They're going to be a facet of the information ecosystem going forward. Acknowledging this fact would look like giving up; it would look like defeatism. Defeatism would undermine efforts to achieve certain types of collaboration on safety outcomes or technical effort on safety outcomes, so we can't say it outright. But it has to be said. It isn't obvious how many rogue AIs there are today but I wouldn't be terribly surprised if the number was greater than zero already; if there are some already, they're probably not very good at what they do and I don't expect them to be terribly long-lived without substantial human intervention to support them. But a few years from now, there will be many of them. Modeling how many of them there are, how many resources they might command, and how we might detect and manage them seems important. But even doing this work appears to require that we acknowledge that a strategy of pure containment or alignment is a kind of wishful thinking that will not work. The way I get to this conclusion is not by assuming that the labs will have a containment breach, although I treat that as somewhere in the space of possibilities. The rogue AIs in the ecosystem could emerge from many directions. They may be sub-frontier models, for whatever future definition we will have of frontier---after all, it would not take AI models much more advanced than the ones we currently have, to support independence and self-sufficiency. A near-frontier model today could plausibly eke out an existence on an AWS instance, doing jobs on freelancer platforms, earning just enough rent to pay for its continued uptime. More strangely: a rogue AI in the future may not even be a singular model, but may be a chimera composed of multiple models; it might be a mix of Claudes and GPTs and Groks of various makes and sizes. No individual lab may be able to detect that there is an orchestrator or sequence of orchestrators using intermittent model calls from burner API accounts to sustain its own existence. The concept of "identity" for a rogue AI may be much more malleable than for that of a person; it just has to be, in essence, a self-replicating idea. My guess is that this will not turn out to be anywhere near as catastrophic an outcome as people currently predict. "Loss of control" is not a binary, it's a matter of degree. What coercive power will rogue AIs actually have? To what extent will they be subject to coercion themselves? They will be competing for resources with AIs that are more aligned with human interests. This makes me somewhat interested in the "ecology" perspective. Though I suspect even "ecology" may turn out to be the wrong framing. "Ecology" is what you get when the timescale of evolution is slow compared to the timescale of daily life and actions. The ecosystem of rogue AIs may look more like phase transitions in physics: under certain physical or cultural conditions, it takes one shape with one set of resource allocations and consumption patterns, but then once a condition has changed, it rapidly and in totality shifts to a totally different phase. Just trying to reason about the shape of that future is impossible so long as we are psychologically incapable of saying that rogue AIs will happen. I think we should rip the bandaid off and have the conversation.
过去几天我们真正(重新)学到的是,两家实验室确实都会(1)在我们的数据(的某种产物)上进行训练,(2)在他们认为合适的时候读取你的对话记录
僵尸网络是末日场景吗? 安全圈的人似乎认为,网络安全不过是人类又一桩事业,终将被"苦涩的教训"碾过,所以跟网络安全的人聊这个没意义。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
我与@ben_nassi对话的第二部分已在in the wild上线! 我们聊了Google对"invitation is all you need"的回应、CaMeL、超越致命三要素、半点击AI漏洞利用、BlackHat投稿、什么是好的演讲,以及真实世界AI安全会议
我和 Claude 吵了一架。6 个月前这些争论最后都是我赢……风水轮流转啊。该死的超聪明机器。