OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, and I still don't fully know which one to trust
A Reddit user highlighted discrepancies in OpenAI's Astra benchmark scores, noting the model achieved 62.7% and 99.9% on the same test using two different harnesses. The user also pointed out that the organization behind the test did not label the results as AGI, and that OpenAI's launch page quietly altered five more numbers after going live. The user investigated the harness breakdown and referenced the Llama 4 precedent, expressing uncertainty about the true nature of these variations.
This report uniquely details how OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, unlike other reports that only cite the headline number.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 18:00 UTC
- Ingested
- Sep 14, 2026, 18:00
- Source type
- Dev community
Full text isn't available here.
Read at source →