RCreddit.com·
暂不在当前实时榜单
Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]
Nonobench is an open benchmark evaluating 49 LLMs on nonogram puzzles, providing row and column clues for models to return the full grid without tools or retries. Solve rates decrease significantly with puzzle size, from 85% for 5x5 to 20% for 15x15. GPT-6 Astra solved all 30 Standard puzzles, while Claude Opus 5.5 solved 8 of 10 Hard mode puzzles, with 11 of 15 models solving none. The benchmark's code is open source under the MIT license.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月4日 13:00 UTC
- 收录
- 2026年10月4日 13:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →