Three agents · One test task
The benchmark refused to end.
Jeff gave the same opening prompt to Claude Code, Grok Build, and Codex to see which coding agent could turn a loose idea into something playable. In that test, Codex was the clear winner.
The comparison should have ended with a result. Instead, the Codex build had enough spark to demand another prompt, then another. What began as a test task between three coding agents became an original, epic game—and a week-long coding session shaped by direct playtests, ambitious new ideas, blunt corrections, and hundreds of deliberate decisions.