跳到正文

Redwood Research

今天还没有收录新动态
10月3日周六
  1. Buck Shlegeris · 收录 · 原文 39

    Redwood Research 与 Anthropic 合作发布概念推理指数 CRI

    Redwood Research 团队与 Anthropic 合作开发了概念推理指数(CRI),用于衡量模型在缺乏廉价可靠反馈的领域中的推理能力,例如判断某项实验能否说明未来远超人类的模型的行为。CRI 的每一条数据都由团队研究员人工核查以保证质量。评测中 0 分对应三项基准上全部随机猜测,100 分为最高分,团队估计真实性能上限为 91。官方排行榜网站为 https://conceptualreasoning.ai/,将持续更新。Buck Shlegeris 转发了 Em 及其团队这项工作并表示期待。

    引用Emery Cooper@emwcooper

    We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?" Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this. Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality. This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below. Official leaderboard website which we'll keep up-to-date: https://conceptualreasoning.ai/

  2. Anthropic Alignment Science · 收录 · 原文 54

    Redwood Research 与 Anthropic 发布概念推理指数 CRI

    Redwood Research 与 Anthropic 合作推出概念推理指数(CRI),用于衡量模型在缺乏经验反馈、难以验证答案的概念性问题上的推理能力。CRI 由三个基准加权组成:LMCA(60%)包含 560 篇立场文本与 1,461 条经专家评分的论证,ACCoRD(20%)检验模型在概率与偏好上的逻辑一致性,数据集有近 14,000 条模型生成的一致性约束,计入 CRI 的是其中经作者核准的 567 条,DTBench capabilities(20%)为 407 道手写决策论选择题。截至 2026 年 8 月 10 日,得分最高的 Opus 5 为 73.6(95% CI ±2.1),低于约 91 的估计上限;自 2024 年底以来分数大致线性上升,未见放缓。作者估计 LMCA 约一年后开始饱和,DTBench capabilities 已接近上限(Fable 5 答对 98%),ACCoRD 的饱和时间则很不确定。

    推荐理由Redwood Research 与 Anthropic 联合发布概念推理基准,给出各实验室模型在无经验反馈任务上的得分与饱和预测。

已经到底了