返回
RCreddit.com

Frontier AI models are beginning to cross the human baseline on SimpleBench

模型发布
时间与来源
发布
09/06 01:58
收录
09/06 14:00
来源类型
开发者社区
档位
社区
信源状态
正常
档位是按信源手工设定的编辑判断,不是逐条打分。

EDIT: GPT-6 Astra hasn't been measured yet. They usually take a couple of days since release. u/Alt_Restorer also pointed out "the benchmark is private, so Phillip only runs it if there's an API that doesn't retain data. I'm not sure whether Astra has that yet."

SimpleBench is designed around reasoning that humans tend to find relatively easy but LLMs have historically struggled with - things like commonsense reasoning, spatial and temporal reasoning, social judgement, and questions designed to catch shallow pattern-matching and/or misleading assumptions.

The plotted model score is average accuracy across 5 runs. Random guessing would score 16.7% - very close to where models started just over two years ago.

The baseline comes from 9 native-English, non-specialist participants, each of whom answered a random subset of 25 questions. Participants had at least high-school-level maths proficiency. So likely above avg.