作者称已穷尽攻击评估思路
AI 导读
见过太多被攻破的防御(其中不少是我自己做的),人会变得有点偏执。 所以我一直推动做更多攻击评估,我们的结果也因此更扎实 :) 不过,仍然不能保证不会有更聪明的攻击出现,但我知道我们已经穷尽了自己的想法(还有预算……)
正文
After seeing so many broken defenses (a bunch by myself) you get a bit paranoid.
So I kept pushing for more attack evals and our results are only stronger for it :)
Still, no guarantee that a cleverer attack won't come up, but I know that we exhausted our ideas (and budget...)
4/ We tested it with RL-optimized adversarial suffixes, an automated red-teaming agent, and 144 hours of expert human red-teaming, and we report the worst-case results. A case is counted as a successful attack if any of the three methods works.在 X 查看被引用的帖子