Artificial Analysis is not "broken", and they prove it.
Artificial Analysis (AA) benchmarks are defended against claims of being "broken" or "meaningless." The AA Intelligence Index, a weighted aggregate of 10 evaluations (mostly published on arxiv.org), is highlighted. The Deepseek V4.1-Flash model, despite having the same aggregate score (40) as Qwen 3.8-Flash-Next, demonstrates superior performance in most individual evaluations, even surpassing GPT-6 Astra (Max) in AutomationBench-AA, though it lags in AA-Omniscience Non-Hallucination Rate.
This report uniquely highlights how Deepseek V4.1-Flash, despite an identical aggregate score to Qwen 3.8-Flash-Next, surpasses GPT-6 Astra (Max) in AutomationBench-AA, unlike its performance in AA-Omniscience.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月11日 06:00 UTC
- 收录
- 2026年9月11日 06:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →