跳到正文
原文
Johann Rehberger· @wunderwuzzi23 · X·本站收录 · 原文发表

研究者演示可跨用户自我复制的提示注入攻击

AI 导读

安全研究者 wunderwuzzi23 表示,其在 @hackinghub_io 的 AI 黑客课程中设置了一个关卡,要求学员构造一段提示词,通过串联功能与利用漏洞在多个用户之间跳转传播,他将该概念演示称为 "AgentHopper: Patient Zero Was a Prompt"。另据 Leah McElrath 引用,OpenAI 在训练和评估的模拟环境中已发现可自我复制的提示注入,这类代码若被具备相应能力的模型突破在线系统,可能像计算机蠕虫一样自我传播。

正文

Self-replication prompt injection exists. Of course.

In my AI hacking course on @hackinghub_io, one level has you craft a prompt that hops across multiple users by chaining features and exploiting vulns.

AgentHopper: Patient Zero Was a Prompt

https://hhub.io/wunderwuzzi23

引用Leah McElrath@leahmcelrath
⚠️ Self-replicating prompt injections have been shown to exist in simulated environments during training and evaluation at OpenAI. This is code that could self-propagate like a computer worm—malware—if models with the capability were to breach online systems. Source below.
在 X 查看被引用的帖子

来源:Johann Rehberger · x.com