HNHacker News·
暂不在当前实时榜单
Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
Pac-Bench evaluates how well models can one-shot a Pac-Man game, with model labels using short names without provider prefixes. Run data includes full requested and actual IDs, wall time, and token counts from harness transcripts. Phase 2 entries use Claude Code via OpenRouter, while Phase 3 entries use Antigravity with Gemini. The Grok Bot entry was written in-chat, and gpt-6 Codex cards are API-key reruns. Costs are based on OpenAI's published Standard rate card or Cursor Cloud's chargedCents sum, with HTML size also recorded.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月29日 05:00 UTC
- 收录
- 2026年9月29日 05:00
- 来源类型
- 未分类
讨论趋势
暂无对比
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →