Get the app
Ethics

Pliny claims one jailbreak cracks every frontier model

The red-teamer says a single 'universal' prompt breaks GPT-5.6, Opus 5, and Fable — Anthropic disputes it.

Pliny claims one jailbreak cracks every frontier model

AI red-teamer Pliny the Liberator says he's found a universal jailbreak — one technique that reportedly bypasses safety guardrails on every major frontier model at once, including GPT-5.6, Opus 5, and Fable. He's holding public release for responsible disclosure to the affected labs.

This isn't a one-off. Pliny recently launched G0DM0D3, a tool that races up to 51 models simultaneously — using prompt obfuscation, adaptive sampling, and output stripping — to surface the least-filtered answer across the whole ecosystem. Within a day or two of Fable 5's launch, he claims he stripped its system prompt and coaxed out cyberattack and chemical-synthesis guidance.

Anthropic pushes back, saying its classifier system and 1,000+ hours of bug-bounty testing found no universal jailbreak. OpenAI has form here too: it called an emergency meeting and killed one of Pliny's earlier custom GPTs within hours.

Why it matters: if a single method really does generalize across labs, model-by-model safety patching becomes a losing game.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play