Get the app
Ethics

Anthropic's Unhackable AI Hacking Tool Got Hacked

Unauthorized actors accessed Mythos — Anthropic's withheld cyberweapon AI — via a breached third-party contractor on the day it launched.

Anthropic's Unhackable AI Hacking Tool Got Hacked

Claude Mythos, Anthropic's most capable model yet, was never meant to go public. It can exploit zero-days across every major OS and browser, succeeding on the first attempt 99% of the time for unpatched vulnerabilities. Anthropic restricted it to a closed defensive-use program instead of a public release.

On April 21 — the same day Mythos launched — an unauthorized group gained access through Mercor, a training contractor. They reverse-engineered Anthropic's URL naming patterns from prior model leaks (Anthropic accidentally exposed Mythos in a misconfigured content store back in March), then used still-active contractor credentials to get in. The group reportedly also accessed other unreleased models in Anthropic's pipeline.

Anthropic confirmed it is "investigating a report of access through one of our third-party vendor environments" but says no core systems were compromised. Notably, CISA — the top US cybersecurity agency — doesn't even have sanctioned access to Mythos.

Why it matters: The most dangerous AI hacking tool ever built was breached on day one through a contractor — the exact supply-chain attack vector it was designed to help defend against.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play