Get the app
LLMs

Fable 5.1 Is Two Days Old and Devs Are One-Shotting GTA Clones

Anthropic's new frontier model doubled its agentic research score and cut cache reads 75% — and X is flooded with one-prompt games.

Two days after launch, the Claude Fable 5.1 timeline is a demo reel: a GTA-style multiplayer clone set in NYC, a Call of Duty knockoff, an F-35A, Super Mario from a single prompt. Viral one-shots are the noisiest signal, but the benchmarks underneath them are the real story.

Fable 5.1 hits 52.6% on Terminal-Bench-Science 0.1 — more than double Fable 5's 24.7%, and well past Opus 5's 29.0%. Agentic coding (Terminal-Bench 4.0) jumps from 42.0% to 55.8%, CursorBench 3.2.0 lands at 73.4%, and Humanity's Last Exam with tools reads 65.0%. The pattern: big gains on long, multi-step work, modest ones on the short tasks the model already nailed.

Pricing is where it gets interesting. Sticker stays at $10/M input, $50/M output, but cache reads dropped 75% — $1.00 to $0.25 per million. Anthropic pegs that at ~25% cheaper on typical workloads and up to 45% on agentic ones. Stateless one-shot calls see none of that discount. Claude Mythos 5.1 is the same model behind trusted-access safeguards for cyber and life sciences.

Why it matters: the cache-read cut makes long-running agents dramatically cheaper to run, which matters more than any game clone on your feed.

Sources

Primary: the company, paper or repository

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play