Skip to content
RCreddit.com·
Not on the current live radar

GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology

AI summary

A benchmark comparing GPT-5.6 Luna and GPT-6 Astra across 50 real PRs from projects like Cal, Sentry, Discourse, Keycloak, and Grafana revealed that Astra found 92 confirmed bugs compared to Luna's 69. However, Luna caught 75% of the bugs at only 3.6% of the cost, with every finding independently verified. The methodology of this benchmark is currently seeking feedback.

Why this one

Unlike many benchmarks, this report uses 50 real-world pull requests from major open-source projects, offering a more practical comparison of GPT-5.6 Luna and GPT-6 Astra.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 14, 2026, 23:01 UTC

Ingested
Sep 14, 2026, 23:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com