跳到正文
RCreddit.com·
暂不在当前实时榜单

GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology

AI 摘要

A benchmark comparing GPT-5.6 Luna and GPT-6 Astra across 50 real PRs from projects like Cal, Sentry, Discourse, Keycloak, and Grafana revealed that Astra found 92 confirmed bugs compared to Luna's 69. However, Luna caught 75% of the bugs at only 3.6% of the cost, with every finding independently verified. The methodology of this benchmark is currently seeking feedback.

为什么是这条

Unlike many benchmarks, this report uses 50 real-world pull requests from major open-source projects, offering a more practical comparison of GPT-5.6 Luna and GPT-6 Astra.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月14日 23:01 UTC

收录
2026年9月14日 23:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com