Skip to content
HNHacker News·
Not on the current live radar

Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?

AI summary

Pac-Bench evaluates how well models can one-shot a Pac-Man game, with model labels using short names without provider prefixes. Run data includes full requested and actual IDs, wall time, and token counts from harness transcripts. Phase 2 entries use Claude Code via OpenRouter, while Phase 3 entries use Antigravity with Gemini. The Grok Bot entry was written in-chat, and gpt-6 Codex cards are API-key reruns. Costs are based on OpenAI's published Standard rate card or Cursor Cloud's chargedCents sum, with HTML size also recorded.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 29, 2026, 05:00 UTC

Ingested
Sep 29, 2026, 05:00
Source type
Unclassified

Full text isn't available here.

Read at source →
Source·Hacker News·jonclegg.github.io