Skip to content
RCreddit.com·
Not on the current live radar

Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode

AI summary

Nonobench v1.2 evaluated 43 LLMs on nonogram puzzles, with GPT-6 Astra achieving a perfect 30/30 score on 15x15 puzzles. The open-weight DeepSeek V4 Pro tied for 4th with 83%, while DeepSeek V4.1 Flash scored 77%. In the new 20x20 Hard mode, Opus 5.5 solved 8/10 puzzles, but no open-weight model managed to solve any, scoring 0/10. Raw data, API, and code are available on nonobench.com and GitHub.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 27, 2026, 17:00 UTC

Ingested
Sep 27, 2026, 17:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com