跳到正文
原文
Owain Evans· @OwainEvans_UK · X·本站收录 · 原文发表

研究者探讨 AI 模型在难以验证的“模糊推理”任务上的能力与评估困境

AI 导读

研究者 Owain Evans 提出,模型在哲学、科研品味、政治或商业策略等“模糊推理”任务上可能不如数学或网络攻击,但进步可能很快。数学领域有 Lean 证明可先验证再阅读,模糊任务缺乏此类机制,且人类对洞察的判断本身分歧更大。他认为模糊任务可能被低估激发,若进展呈锯齿状则不易察觉,需警惕。

正文

Some fuzzy ideas about fuzzy hard-to-verify reasoning in AI models. Fuzzy reasoning = conceptual/philosophical reasoning, research taste in science, strategy in politics or business, etc.

It’s plausible models are not as good at fuzzy reasoning as they at math or cyber/hacking, but also that they are improving fast. Pretraining does not favor verifiable over fuzzy tasks and pretrains are improving.

In math, when models have new ideas, it seems they are often bad at explaining them to humans. e.g. they produce the “worst proof write-up ever” and a bunch of work is needed by humans and models to make it readable. My guess is that models are worse at explaining things that humans didn’t already explain before. *So if models had impressive fuzzy insights, they would probably be also be bad at explaining them today.*

With math, we can have a Lean proof and so only bother reading proofs for verified claims. We don’t have that for fuzzy tasks.

For some fuzzy tasks, humans are worse at judging insights and it’s inherently harder. E.g. Mathematicians agree on proofs being correct (even without verification) but we have much less agreement about philosophy insights (even with much human-time to evaluate them). Even in AI research, there’s less agreement on what counts as an insights in advance of experiments.
(Sidenote: various people had prescient insights about AI years in advance but they didn’t gain much traction at the time!)

Overall, it’s plausible that fuzzy tasks are under-elicited in models and so there could be more rapid progress with some tricks for better elicitation (e.g. some kind of self-play or distillation tricks). Progress in fuzzy tasks may not be as obvious as verifiable ones, and so we should watch out for that. This is especially true if the fuzzy task progress is jagged. (Progress in AI research taste should be easier to recognize than in philosophy because of quick experimentation in the former case.)

P.S. Recall Wittgenstein on heliocentrism. What would it look like to be in a world where AIs have (jagged) insights in areas that are particularly hard to evaluate?

来源:Owain Evans · x.com