The Pac-Bench project evaluates how effectively various models can recreate a Pac-Man game based on a single prompt: 'Create a Pac-Man game in a single html page.' The project provides model labels without provider prefixes, dates, or effort suffixes. Detailed data on requested and actual IDs is available under Run data. When two cards share a label, it is noted if a card requested a different model. Performance statistics include wall time and token counts from harness transcripts when available. Phase 2 entries utilize Claude Code directed at OpenRouter, while Phase 3 entries employ Antigravity with Gemini. The Grok Bot entry was generated in-chat from the same prompt without prior access to the task. The gpt-6 Codex cards are reruns using API keys, with costs based on OpenAI's published Standard rate card applied to the captured response usage. Cursor Cloud costs are indicated where applicable, based on the dashboard chargedCents sum. The HTML size represents the entry file as served, with a dash indicating that the run did not record that number.
✓ No loaded language, vague sourcing, or framing detected.
Pac-Bench Tests Model Performance in Creating Pac-Man Game from Single Prompt
The Pac-Bench project assesses the ability of various models to generate a Pac-Man game from a single prompt. It includes performance statistics and details about the models used in the testing phases.
No note attached
on this article.
Original vs. Neutral
Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
Pac-Bench Tests Model Performance in Creating Pac-Man Game from Single Prompt