Artificial Analysis is not "broken", and they prove it.
Artificial Analysis (AA) benchmarks are defended against claims of being "broken" or "meaningless." The AA Intelligence Index, a weighted aggregate of 10 evaluations (mostly published on arxiv.org), is highlighted. The Deepseek V4.1-Flash model, despite having the same aggregate score (40) as Qwen 3.8-Flash-Next, demonstrates superior performance in most individual evaluations, even surpassing GPT-6 Astra (Max) in AutomationBench-AA, though it lags in AA-Omniscience Non-Hallucination Rate.
This report uniquely highlights how Deepseek V4.1-Flash, despite an identical aggregate score to Qwen 3.8-Flash-Next, surpasses GPT-6 Astra (Max) in AutomationBench-AA, unlike its performance in AA-Omniscience.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 11, 2026, 06:00 UTC
- Ingested
- Sep 11, 2026, 06:00
- Source type
- Dev community
Full text isn't available here.
Read at source →