Anthropic 用 Claude 自主推进可扩展监督研究,以 1.8 万美元算力超过人类研究者
AI 导读
Anthropic 公布一项新研究结果:用 Claude 自主推进可扩展监督研究,以性能差距恢复(PGR)作为衡量指标。Claude 在多种不同技术上反复迭代,最终以 1.8 万美元的额度显著超过人类研究者。
正文
New research result: we use Claude to make fully autonomous progress on scalable oversight research, as measured by performance gap recovered (PGR).
Claude iterates on a number of different techniques and ends up significantly outperforming human researchers for $18k in credits.