Astra is a quiet force. Artificial Analysis has updated their benchmark twice in 4 days to reflect its real strength
Astra is proving to be a significant force, leading Artificial Analysis to update its benchmark twice in four days to accurately reflect its capabilities. A similar situation occurred with Sol, which prompted Arena to restructure its coding benchmark, add a full-stack benchmark, and update its webdev benchmark to better reflect real-world performance after Astra's release. This suggests that OpenAI may not be fully optimizing for benchmarks.
Time & source
- Published
- 09/08, 16:36 UTC+0
- Ingested
- 09/09, 15:00 UTC+0
- Source type
- Dev community
- Tier
- Community
- Source status
- Sync delayed
Tier is a per-source editorial setting, not a per-item score.
Something similar happened with Sol and Arena had to restructure their coding benchmark and added a full-stack bench, plus updated the webdev bench to reflect real world performance after Astra released. Looks OpenAI is only lab not benchmaxxing