Three agent harnesses, one game
Claude Code, prime-agent and pi each got the same spec, the same empty checkout and the same model, Opus 5.5 at high thinking, and built a three.js Brick Breaker in one session. Same model, so what differs is the harness around it.
The numbers
Where the tokens went
cache read
cache write
uncached input
output
What the dollars paid for
Play each build
Each game is the build exactly as its harness left it. A or ← and D or → move the paddle; Space launches the ball.
How it was run
- Each builder started in its own copy of
seed/, outside this project, so it never saw the hidden checks or another build. Each got the same prompt: the spec inspec.md, with the instruction to passbash test/gate.sh. - Model:
claude-opus-5-5, thinking high, one session each, no repairs, a 45-minute limit. Each harness ran as installed on the machine, with its own system prompt, tools and context files. - Claude Code ran
claude -pwith Read, Edit, Write and Bash allowed and no MCP servers; prime-agent and pi ran with-p --no-sessionand their default tools. - Tokens and dollars are what each harness reported for the session. Claude Code reports its own cost; prime-agent and pi report a cost per message. All three match Opus 5.5's list prices to the cent.
- After each build the harness ran the builder's own suite, then 13 hidden checks that drive the game through its WebMCP tools (
hidden/checks.py). The builder never sees them. - Saying only "Hello" (Opus 5.5 high, three times each, 9 October 2026) costs Claude Code about 29.3K tokens, prime-agent 26.5K and pi 18.2K: that is each harness's own system prompt and tools before any work.