Get the app
LLMs

Fable 5.1's party trick: tiny 3D dioramas Anthropic never pitched

Anthropic sold Fable 5.1 on agentic coding benchmarks. Devs are grading it on little one-shot Three.js dioramas instead.

Anthropic shipped Claude Fable 5.1 on 1 September with a benchmark sheet aimed squarely at agents: Terminal-Bench-Science jumps from 24.7% to 52.6%, Terminal-Bench 4.0 coding goes 42.0% → 55.8%, and OSWorld computer use lands at 41.7%. Headline pricing is unchanged at $10/$50 per million in/out, but cache reads fall 75% to $0.25/M — which Anthropic translates to roughly 25% cheaper for typical workloads and up to 45% for heavily agentic runs.

None of that is what people are actually posting. The vibe test doing the rounds is neat little dioramas — one-shot Three.js scenes, isometric rooms, walkable toy worlds — a habit inherited from Fable 5, which developers used to spin up dozens of browser 3D worlds, many landing on the first pass. Anthropic's own launch post claims nothing about visual or 3D quality.

That gap is the actual story. The lab measures long-horizon agent work; the community measures whether the model can render a convincing miniature with no assets, no references and no second try. The second test is unscientific and instantly legible — which is exactly why it travels further than the eval table.

Why it matters: model launches now get judged by whatever demo screenshots well, not by the benchmark the lab optimized for.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play