OpenAI's GPT-6 Astra on ARC-AGI-3
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
OpenAI model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.
- Basis
- Running about 2.0× the median of this source's recent listed items
- Triggering item
- OpenAI's GPT-6 Astra on ARC-AGI-3
- Metric comparison
- 108 vs median 54 (20 baseline samples)
- Detected
- 09/03, 22:01
OpenAI's GPT-6 Astra was evaluated on ARC-AGI-3, demonstrating its performance with context-management features. The Provider Adapter harness allows Astra to preserve opaque reasoning state between requests and uses compaction for longer conversations, enabling the model to reuse prior work. This approach helps manage conversations and maintain context effectively, with performance metrics showing varying percentages across different reasoning effort levels, such as 62.7% for max effort and 17.5% for low effort.
Published
03 Sep 2026
Summary
- GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment. , and 99.9% for $19K with a The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. .
- GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.
- A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions.
ARC-AGI-3
ARC-AGI-3 is a benchmark for studying agentic intelligence through novel, abstract, turn-based environments. Agents must explore, infer goals, and build internal models of environments to effectively plan actions without explicit instructions. You can play ARC-AGI-3 yourself .
Your browser does not support embedded video.
These environments only contain core knowledge priors and are difficulty-calibrated through controlled testing with human participants. Humans can solve 100% of the environments .
The goal of the ARC-AGI series is to measure the “residual gap” between current artificial intelligence and AGI. We define AGI as a system’s ability to acquire any skill a human can, as efficiently as a human can.
ARC-AGI-3 is the third generation of the ARC-AGI benchmark series . It tests agentic capabilities beyond ARC-AGI-1 and ARC-AGI-2 . Each generation expands on the one before it - as frontier AI capabilities advance, our benchmarks must advance with them.
ARC-AGI-3 tests four components of agentic intelligence:
- Exploration: In real-world environments, information is rarely provided passively. Agents must actively obtain it by interacting with their surroundings.
- Modeling: Agents must turn raw observations into a generalizable model that can predict future states and outcomes.
- Goal-setting: Agents must identify target future states with only sparse rewards.
- Planning and execution: Agents must map a path from their current state to a goal, course correcting as new information appears.
Astra Results
GPT-6 Astra achieves state-of-the-art scores on ARC-AGI-3 with both the Standard and Provider Adapter harnesses. Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing the total number of model calls and tokens. View the full results . With our Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment. , OpenAI’s Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K. With the The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. , Astra (high) scores 99.9% for $19K. Both are state-of-the-art scores. See the full leaderboard .
At max reasoning effort, Astra solves games more efficiently, requiring fewer actions and therefore lowering total cost relative to the other reasoning-effort levels.
Reasoning effort Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment. The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. max 62.7%, $26,098 98.6%, $17,332 xhigh 59.3%, $37,317 98.4%, $18,147 high 54.8%, $40,705 99.9%, $18,817 medium 38.6%, $48,090 98.4%, $19,285 low 17.5%, $38,166 98.0%, $21,298 none 35.2%, $49,791 96.7%, $23,457
For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted. 1
Analysis