跳到正文
原文
Alexander Panfilov· @kotekjedi_ml · X·本站收录 · 原文发表

蒸馏攻击防御评估新发现

AI 导读

但事实证明,你可能根本不需要原始推理轨迹…… 来看看这项有趣的工作,我们比较了基于原始轨迹的蒸馏与其他"推理替代物"(包括摘要)的效果。

正文

But it turns out you may not need the raw reasoning traces at all...

Check out this cool work where we compare distillation on raw traces against other "reasoning surrogates" (including summaries).

引用Shidan Javaheri@shidan_javaheri
Excited to share our recent work! All existing defenses against distillation attacks are evaluated after distillation, implicitly assuming no further RL training. We find this makes existing evals give a false sense of security, and that RL makes simple attacks effective
在 X 查看被引用的帖子

来源:Alexander Panfilov · x.com