RCreddit.com·
暂不在当前实时榜单
GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology
A benchmark comparing GPT-5.6 Luna and GPT-6 Astra across 50 real PRs from projects like Cal, Sentry, Discourse, Keycloak, and Grafana revealed that Astra found 92 confirmed bugs compared to Luna's 69. However, Luna caught 75% of the bugs at only 3.6% of the cost, with every finding independently verified. The methodology of this benchmark is currently seeking feedback.
Unlike many benchmarks, this report uses 50 real-world pull requests from major open-source projects, offering a more practical comparison of GPT-5.6 Luna and GPT-6 Astra.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月14日 23:01 UTC
- 收录
- 2026年9月14日 23:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →