Simon Willison tests Claude Fable 5.1's pelican animation benchmark across all five reasoning levels, finding some puzzling results where low and medium effort modes produced no visible reasoning traces despite high output token counts. Anthropic's Fable 5.1 also debuted strong scores on the new Terminal-Bench-Science 0.1 benchmark at 52.6%, outpacing Opus 5 and GPT-5.6 Sol.
Background
Claude Fable 5.1 is Anthropic's latest AI model release, positioning itself strongly in coding and scientific research benchmarks. Willison has maintained a whimsical but increasingly popular 'pelican benchmark' as an informal measure of model creative and visual capabilities.
- Source
- Simon Willison
- Published
- Sep 2, 2026 at 07:57 AM
- Score
- 5.0 / 10