跳到正文
RCreddit.com·
暂不在当前实时榜单

Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode

AI 摘要

Nonobench v1.2 evaluated 43 LLMs on nonogram puzzles, with GPT-6 Astra achieving a perfect 30/30 score on 15x15 puzzles. The open-weight DeepSeek V4 Pro tied for 4th with 83%, while DeepSeek V4.1 Flash scored 77%. In the new 20x20 Hard mode, Opus 5.5 solved 8/10 puzzles, but no open-weight model managed to solve any, scoring 0/10. Raw data, API, and code are available on nonobench.com and GitHub.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月27日 17:00 UTC

收录
2026年9月27日 17:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com