GPT-6 Astra is the first model to solve all 30 puzzles in my nonogram benchmark
GPT-6 Astra is the first model to solve all 30 puzzles in a nonogram benchmark, which has been evaluating LLMs since January. In the new Hard mode, consisting of ten random 20x20 puzzles, Astra solved 5 out of 10. It successfully completed all five puzzles solvable line by line but failed on those requiring deeper search. Claude Opus 5.5 currently leads this mode with 8 out of 10 puzzles solved, while 11 of 14 models couldn't solve any of the 20x20s. The benchmark is public and open source at nonobench.com.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 29, 2026, 00:00 UTC
- Ingested
- Sep 29, 2026, 00:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →