Back
RCreddit.com

Frontier AI models are beginning to cross the human baseline on SimpleBench

Model release
Time & source
Published
09/06, 01:58
Ingested
09/06, 14:00
Source type
Dev community
Tier
Community
Source status
Healthy
Tier is a per-source editorial setting, not a per-item score.

EDIT: GPT-6 Astra hasn't been measured yet. They usually take a couple of days since release. u/Alt_Restorer also pointed out "the benchmark is private, so Phillip only runs it if there's an API that doesn't retain data. I'm not sure whether Astra has that yet."

SimpleBench is designed around reasoning that humans tend to find relatively easy but LLMs have historically struggled with - things like commonsense reasoning, spatial and temporal reasoning, social judgement, and questions designed to catch shallow pattern-matching and/or misleading assumptions.

The plotted model score is average accuracy across 5 runs. Random guessing would score 16.7% - very close to where models started just over two years ago.

The baseline comes from 9 native-English, non-specialist participants, each of whom answered a random subset of 25 questions. Participants had at least high-school-level maths proficiency. So likely above avg.