RCreddit.com·
暂不在当前实时榜单
GoBench: Evaluating LLMs on the game of Go [R]
GoBench evaluates LLMs on 9x9 Go games against KataGo opponents, measuring general reasoning ability and showing a strong correlation (r=0.83) with ARC-AGI 2. GPT-6 Astra max achieved 2500 Elo, while Codex with Astra, using coding tools and two hours of preparation, reached 3560 Elo. This is still lower than the best KataGo's 4400 Elo, indicating the leaderboard remains highly unsaturated. The project provides code and a paper for further details.
This evaluation uniquely correlates LLM Go performance with ARC-AGI 2 (r=0.83), unlike other Go benchmarks, and shows GPT-6 Astra max reaching 2500 Elo, far below KataGo's 4400 Elo.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月16日 22:00 UTC
- 收录
- 2026年9月16日 22:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →