RCreddit.com·
暂不在当前实时榜单
Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode
Nonobench v1.2 evaluated 43 LLMs on nonogram puzzles, with GPT-6 Astra achieving a perfect 30/30 score on 15x15 puzzles. The open-weight DeepSeek V4 Pro tied for 4th with 83%, while DeepSeek V4.1 Flash scored 77%. In the new 20x20 Hard mode, Opus 5.5 solved 8/10 puzzles, but no open-weight model managed to solve any, scoring 0/10. Raw data, API, and code are available on nonobench.com and GitHub.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月27日 17:00 UTC
- 收录
- 2026年9月27日 17:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →