Skip to content
RCreddit.com·
Not on the current live radar

Nonobench: an open benchmark of 49 LLMs on nonogram puzzles, public and open source [P]

AI summary

Nonobench is an open benchmark evaluating 49 LLMs on nonogram puzzles, providing row and column clues for models to return the full grid without tools or retries. Solve rates decrease significantly with puzzle size, from 85% for 5x5 to 20% for 15x15. GPT-6 Astra solved all 30 Standard puzzles, while Claude Opus 5.5 solved 8 of 10 Hard mode puzzles, with 11 of 15 models solving none. The benchmark's code is open source under the MIT license.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 4, 2026, 13:00 UTC

Ingested
Oct 4, 2026, 13:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com