GPT-6 Astra vs Claude Fable 5.1 based on the benchmarks available so far
- 发布
- 09/04 03:26
- 首次发现
- 09/04 17:00
- 类型
- 开发者社区 · RSS
目前已收集并分析了 GPT-6 Astra 和 Claude Fable 5.1 的公开基准测试结果。根据 Artificial Analysis 的数据,Fable 5.1 在 Intelligence Index 和 Coding Agent Index 方面表现更优。然而,Coding Agent Index 的比较受到外部环境的影响,因此并非纯粹的模型对模型测试。期待更多使用过这两个模型的用户提供实际使用对比。
I collected the public benchmark results I could find for GPT-6 Astra and Claude Fable 5.1 and put them side by side.
GPT-6 Astra leads on several published coding, science and agent benchmarks, including Terminal-Bench and DeepSWE, while Fable 5.1 still performs better on some independent evaluations.
For example, Artificial Analysis currently gives Fable 5.1 a higher Intelligence Index and Coding Agent Index, although the coding-agent comparison also depends on the surrounding harness, so it isn't a pure model-to-model test.
Another interesting point is that both models have the same headline API pricing at $10/M input and $50/M output, although caching and long-context pricing differ.
I don't think there is enough independent data yet to say one is clearly better overall, but the current results make for an interesting comparison.
I collected the numbers here: https://llmlearner.com/compare/gpt-6-astra-vs-claude-fable-5-1
Would be interested to see more real-world comparisons from people who have used both.