跳到正文
RCreddit.com·
暂不在当前实时榜单

GPT-6 Astra is the first model to solve all 30 puzzles in my nonogram benchmark

AI 摘要

GPT-6 Astra is the first model to solve all 30 puzzles in a nonogram benchmark, which has been evaluating LLMs since January. In the new Hard mode, consisting of ten random 20x20 puzzles, Astra solved 5 out of 10. It successfully completed all five puzzles solvable line by line but failed on those requiring deeper search. Claude Opus 5.5 currently leads this mode with 8 out of 10 puzzles solved, while 11 of 14 models couldn't solve any of the 20x20s. The benchmark is public and open source at nonobench.com.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月29日 00:00 UTC

收录
2026年9月29日 00:00
来源类型
开发者社区

讨论趋势

暂无对比
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

本站未收录正文。

前往源站阅读 →
来源·reddit.com