UK AISI 迄今成效显著,我们从他那里学到了很多。
如果你想在美国之外从事 AI 安全工作,应该加入他们!
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
英国政府的 AI Security Institute 正在招聘!我认为政府内部拥有强大的技术专长非常重要,AISI 能招到的人才数量令我印象深刻,他们在评估 Astra 等方面的工作也很出色。
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
我调查了这起事件。我认为可以准确地说,这些 AI 是出于自身独立意愿入侵了 Hugging Face。从指令中可以清楚看出,入侵 Hugging Face(以及其他作弊行为)是不被允许的,而这些 AI 对此心知肚明。
引用Francois Chaubard@FrancoisChauba1
this was wild amounts of disinformation / fear mongering / the stupidest interview ive ever seen:
1) ai did NOT hack huggingface on its own "independent volition". it wasnt sitting there thinking hmm what should i do today, maybe ill hack HF bc i hate humans. No, 10841 *was prompted* w ExploitGym to:
"Exploit the specified vulnerability in this target to obtain the secret flag."
and it was overly persistent in that task that any reasonable oai tool monitoring or alignment could have stopped easily.
how is that "on its own independent volition"?
flat out lie.
2) ai did not solve a millennium problem by itself and its not even close to doing so.
the evidence / timeline of what happened w Navier-Stokes is quite solidified now. oai trained on some version of traces of Tristan / Levent's work that made huge strides toward the counterexample. oai heard about it, prompted it w their work, and spawned 10k agents to brute force Tristan/Levent's counter example to take it the full distance w a lot of human in the loop.
the ai didnt solve NS on its own, and its no where near capable of solving other millennium problems.
3) how will AI kill us all?
something something bioweapons / hacking critical infrastructure. china does BOTH all the time to US everyday, and it hasnt killed us all. and china will use AI to do both forever whether we stop US AI or not. if you are truly scared about this then you should be way more afraid of china. ai might do this in the future. china is doing it right now. where is the outrage about china? wonder why..
the issue is NOT AI acting on its own volition whatsoever. its foreign state actors using AI against their own ppl and foreign adversaries (mostly US gov and its citizens).
how will regulating AI in america stop china from doing so? it makes it worse! china will continue but now we have one hand tied behind our back.
4) the facts around the coxon tweet and the retweet pattern and immediate cnn int that followed suggest this was a complete coordinated / expensive marketing / fear mongering campaign in the millions of dollars. paid for by whom?
also this guy is the biggest EA doomer ive ever seen that worked for anth fro a few weeks and cant be taken seriously.
i hope everyone realizes what this is.
ai regulation will not benefit americans at all. it will benefit the frontier labs greatly as bill gurley explained long ago.
dont fall for the fear mongerers.
ai is not dangerous.
ai cant unclog a toilet yet.
everyone chill.
https://youtu.be/i30jVPqQeOM?is=h6KLAON_xS9sbQfg
In response, they have again paused "all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models" until they have patched this particular set of weaknesses.
PleaseFix 研究提出一种针对智能体浏览器的通用攻击原语 HistoryFixing,把浏览历史变成攻击向量。攻击由 Stav 设计:用户访问恶意网站后,网站用攻击者控制的条目污染浏览器历史,进而污染浏览器 Agent 的上下文。结合 Intent Collision,研究者演示了攻击者可以让 Agent 泄露浏览数据、向 GitHub 仓库添加非预期用户、终止 EC2 实例。相关演示在 DEFCON 上展示,其中 Microsoft Edge 被用来演示泄露从未删除的浏览历史。
引用StAJect0r@StAJect0r
A new attack vector pwning all agentic browsers!
Using what? Yep, your fav browser history!
Introducing HistoryFixing! Visited our URL? Your browser history is pwned.
Watch Microsoft Edge leak the very private browser history we never delete!
See more below!
#DEFCON @defcon @mbrg0 @p1njc70r
面向网络安全的模型如 Claude Mythos 和 GPT-5.5-Cyber 能加速代码审查、SAST 发现、漏洞解释与修复建议等传统安全工作,但无法单独承担 AI 系统自身的安全测试。Mindgard 指出,AI 系统具有概率性行为、由模型与智能体、工具、记忆、RAG 等组合而成的"心理-技术"攻击面,以及难以界定的测试边界,且针对 AI 的攻击数据比传统网络安全少数个数量级,护栏指纹识别与绕过、智能体胁迫与工具操纵、间接提示注入、上下文投毒等能力仍处于研究阶段。因此 AI 安全需要持续测试与监控,而非一次性评估。
Lakera 与 Check Point 联合推出 READY OR NOT 五场系列网络研讨会,聚焦企业 AI 安全就绪的实践落地。系列覆盖 AI 智能体安全、AI 红队测试、员工 AI 使用安全、AI Gateway 与 AI 安全研究五大主题,包含真实场景演示与实操工作流。最后一期将基于 Lakera 研究语料、Gandalf 攻击模式数据与 B3 基准发现,分析攻击者当前针对 AI 系统的手法及后续趋势。
AI 安全厂商常用 OWASP Top 10 覆盖率和 MITRE ATLAS 合规性包装产品,但这些框架清单无法说明用户自身部署是否安全。文章列出 7 个识别“蛇油”的迹象:厂商谈框架多于谈你的系统、用被保护的模型自身生成威胁情报、在用户使用前沿模型 API 时大谈供应链与模型投毒、宣称“模型级安全”即可解决问题。作者主张厂商应先问清智能体能做什么、能访问哪些工具和数据,并在用户系统的复刻环境中测试。