研究者演示可跨用户自我复制的提示注入攻击
AI 导读
安全研究者 wunderwuzzi23 表示,其在 @hackinghub_io 的 AI 黑客课程中设置了一个关卡,要求学员构造一段提示词,通过串联功能与利用漏洞在多个用户之间跳转传播,他将该概念演示称为 "AgentHopper: Patient Zero Was a Prompt"。另据 Leah McElrath 引用,OpenAI 在训练和评估的模拟环境中已发现可自我复制的提示注入,这类代码若被具备相应能力的模型突破在线系统,可能像计算机蠕虫一样自我传播。
正文
Self-replication prompt injection exists. Of course.
In my AI hacking course on @hackinghub_io, one level has you craft a prompt that hops across multiple users by chaining features and exploiting vulns.
AgentHopper: Patient Zero Was a Prompt
https://hhub.io/wunderwuzzi23
⚠️ Self-replicating prompt injections have been shown to exist in simulated environments during training and evaluation at OpenAI.
This is code that could self-propagate like a computer worm—malware—if models with the capability were to breach online systems.
Source below.在 X 查看被引用的帖子