Get the app
LLMs

Fable 5.1 Is Two Days Old and Devs Are One-Shotting GTA Clones

Anthropic's new frontier model doubled its agentic research score and cut cache reads 75% — and X is flooded with one-prompt games.

Two days after launch, the Claude Fable 5.1 timeline is a demo reel: a GTA-style multiplayer clone set in NYC, a Call of Duty knockoff, an F-35A, Super Mario from a single prompt. Viral one-shots are the noisiest signal, but the benchmarks underneath them are the real story.

Fable 5.1 hits 52.6% on Terminal-Bench-Science 0.1 — more than double Fable 5's 24.7%, and well past Opus 5's 29.0%. Agentic coding (Terminal-Bench 4.0) jumps from 42.0% to 55.8%, CursorBench 3.2.0 lands at 73.4%, and Humanity's Last Exam with tools reads 65.0%. The pattern: big gains on long, multi-step work, modest ones on the short tasks the model already nailed.

Pricing is where it gets interesting. Sticker stays at $10/M input, $50/M output, but cache reads dropped 75% — $1.00 to $0.25 per million. Anthropic pegs that at ~25% cheaper on typical workloads and up to 45% on agentic ones. Stateless one-shot calls see none of that discount. Claude Mythos 5.1 is the same model behind trusted-access safeguards for cyber and life sciences.

Why it matters: the cache-read cut makes long-running agents dramatically cheaper to run, which matters more than any game clone on your feed.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play