Simon Willison compares Claude Fable 5 and Codex with GPT-5.6 Sol Ultra on a one-shot game generation task, finding the latter produced a more complex 'Raccoon Heist' game with multi-agent capabilities. The generated version contained a visual bug with oversized eyeballs that required iterative prompting to fix, highlighting current limitations in autonomous code review.
Background
Simon Willison frequently benchmarks AI coding agents by having them generate games from prompts. GPT-5.6 Sol Ultra is a mode that aggressively utilizes sub-agents for complex tasks.
- Source
- Simon Willison
- Published
- Aug 8, 2026 at 03:18 AM
- Score
- 5.0 / 10