返回
RCreddit.com
WARM

The prevalent problem of misleading benchmark reporting (re: Astra)

查看原文
ClaudeOpenAI
时间与来源
发布
09/03 21:23
首次发现
09/04 20:00
类型
开发者社区 · RSS
AI 摘要

OpenAI 报告的 Astra 在 ARC-AGI-3 上的基准测试结果因具有误导性而受到批评。尽管一张截图显示 Astra 取得了 98.6% 的成绩,而 GPT 5.6 Sol 为 7.8%,Claude Opus 5 为 30.2%,但这一结果是在特定提供商适配器工具下实现的。在标准的 ARC-AGI-3 工具下,Astra 的成绩是 62.7%,而 Opus 5 为 30.2%,Sol 为 7.8%。尽管差距仍然显著,但远没有之前那么夸张,这表明 OpenAI 可能存在故意误报的行为。

OpenAI's reported benchmarks for Astra's ARC-AGI-3 is one of the most egregious recent examples I have seen of technically true metric reporting being used to deliberately mislead the masses. For context, there is an OpenAI screencap currently at the top of r/singularity 's hot page of Astra achieving 98.6% on ARC-AGI-3 compared to 7.8% for GPT 5.6 Sol and 30.2% for Claude Opus 5. Holy shit, right? ASI achieved, right?

Unfortunately, those figures taken in a vacuum leave out very important context: Astra's agentic harness had significant additional features that GPT 5.6 Sol and Claude Opus 5 did not have access to - specifically reasoning trace retention and custom compaction (source: https://arcprize.org/leaderboard ).

My main takeaway is basically: The most honest way to compare Astra with Opus 5/Sol on this benchmark would have been to either 1) measure their ARC-AGI-3 performances on the same provider adapter harness (where Astra's 98.6% came from), or 2) compare them on the standard ARC-AGI-3 harness. On the standard harness Astra achieves 62.7% vs Opus 5's 30.2% vs Sol's 7.8%. Still a very large gap, but much less misleading than the comparison OpenAI chose to report. (source: https://arcprize.org/leaderboard)

Not an Anthropic fanboy in any sense of the word, btw. I thought Opus 5 was benchmaxxed and pray on Anthropic's downfall every day. But the Astra benchmark glazing made it clear that restraint needs to be had in people's reactions to its benchmarks (if Opus 5 didn't already convince you to not treat benchmarks as gospel) before anyone has even had time to extensively test it in real world use cases.