Neel Nanda:可解释性尚不足以被依赖
AI 导读
我坚持这一观点——可解释性可能意义重大,也确实足够有用,但远未达到任何人应当依赖我们来确保一切顺利的质量和可靠性水平。
正文
I stand by this line - interpretability MIGHT be a big deal, and is certainly good enough to be useful, but is nowhere near the level of quality and reliability where anyone should be relying on us to ensure things go well
@NeelNanda5 is widely regarded as one of the top two experts on mechanistic interpretability in the world.
“Speaking as an interpretability expert, please do not rely on us to save you on the current trajectory."
https://youtu.be/J38ot52b2-E在 X 查看被引用的帖子