Anthropic's Unhackable AI Hacking Tool Got Hacked
Unauthorized actors accessed Mythos — Anthropic's withheld cyberweapon AI — via a breached third-party contractor on the day it launched.

Claude Mythos, Anthropic's most capable model yet, was never meant to go public. It can exploit zero-days across every major OS and browser, succeeding on the first attempt 99% of the time for unpatched vulnerabilities. Anthropic restricted it to a closed defensive-use program instead of a public release.
On April 21 — the same day Mythos launched — an unauthorized group gained access through Mercor, a training contractor. They reverse-engineered Anthropic's URL naming patterns from prior model leaks (Anthropic accidentally exposed Mythos in a misconfigured content store back in March), then used still-active contractor credentials to get in. The group reportedly also accessed other unreleased models in Anthropic's pipeline.
Anthropic confirmed it is "investigating a report of access through one of our third-party vendor environments" but says no core systems were compromised. Notably, CISA — the top US cybersecurity agency — doesn't even have sanctioned access to Mythos.
Why it matters: The most dangerous AI hacking tool ever built was breached on day one through a contractor — the exact supply-chain attack vector it was designed to help defend against.
Sources
- Unauthorized group has gained access to Anthropic's Mythos, report claims techcrunch.com
- Anthropic withholds Mythos Preview model because its hacking is too powerful axios.com
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.