@NeelNanda5 is widely regarded as one of the top two experts on mechanistic interpretability in the world.
“Speaking as an interpretability expert, please do not rely on us to save you on the current trajectory."
https://youtu.be/J38ot52b2-E
Are SAEs dead?
Will they save us from neuralese?
Should I just use a probe?
We get these questions all the time. Part 2 of our educational series on applied interpretability explains what SAEs are good for, when *not* to use them, and what to use instead. 🧵
Joe Carlsmith 在系列文章第四篇中提出"AI for AI safety",即用前沿 AI 劳动强化安全进展、风险评估与能力约束三大安全因素,以让 AI 安全反馈回路追上或约束 AI 能力反馈回路。他提出"AI for AI safety 甜蜜点"概念,指前沿系统足以大幅改善安全因素、但尚不足以在现有对策下剥夺人类权力的能力区间,并指出该窗口未必存在且难以持久。文章最后列出最严重的担忧,包括诱导/评估失败、差异性破坏与危险的失控选项,以及时间与政治意愿等现实约束,并将在下一篇聚焦自动化对齐研究。