研究者披露 API 漏洞可读取多家前沿模型的加密思维链
Jonas Geiping 等人发布报告,披露一处 API 漏洞可读取多家前沿模型的加密思维链,并验证其推理 token 数与 API 计费的 thinking token 在多数提示下为 1:1。作者指出,公开分享的推理轨迹存在直接隐私风险,团队能解码网上许多 JSON 轨迹并发现其中的私密信息。调查期间他们还因泄露的生产密钥获得了 HuggingFace 的访问入口,但仅执行了 whoami 未进一步使用。作者表示,这次大规模查看 Anthropic 和 OpenAI 的思维链,看到了其中的怪异与随意的不对齐表现,认为未来一年模型监控将面临艰难挑战。他同时主张向所有用户开放思维链,以实现更广泛、更多元的监督。
披露了通过 API 漏洞读取前沿模型加密思维链的方法,并指出公开推理轨迹中的隐私泄露风险。
Earlier today we release our report about a vulnerability that allowed us to read out the encrypted thinking traces from many frontier models (thread below!):
A few thoughts:
First, there is an immediate privacy concern with publicly posted reasoning traces (which is also why we took time to release the report after the initial disclosure). We were able to decode the thinking of many json traces posted online, and found private info in there.
Ironically, during this investigation we also had entry into HF during the cybersec incident due to a leaked prod key (but did not exercise the key beyond a whoami ;)).
Another big update from this for me was actually seeing thinking traces at scale from Anthropic and OpenAI, and all the weirdness and 'casual' misalignment they contain (examples below). I do think model monitoring has an uphill battle ahead of us in the coming year.
---
Finally, there is something to be said for just making thinking traces accessible to all users. I do believe that we would achieve a much broader, much more pluralistic form of oversight through broadly accessible thinking, that would make model deployment safer.
We can finally talk about it:
We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.
We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.在 X 查看被引用的帖子