GoBench: Evaluating LLMs on the game of Go [R]
GoBench evaluates LLMs on 9x9 Go games against KataGo opponents, measuring general reasoning ability and showing a strong correlation (r=0.83) with ARC-AGI 2. GPT-6 Astra max achieved 2500 Elo, while Codex with Astra, using coding tools and two hours of preparation, reached 3560 Elo. This is still lower than the best KataGo's 4400 Elo, indicating the leaderboard remains highly unsaturated. The project provides code and a paper for further details.
This evaluation uniquely correlates LLM Go performance with ARC-AGI 2 (r=0.83), unlike other Go benchmarks, and shows GPT-6 Astra max reaching 2500 Elo, far below KataGo's 4400 Elo.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 16, 2026, 22:00 UTC
- Ingested
- Sep 16, 2026, 22:00
- Source type
- Dev community
Full text isn't available here.
Read at source →