Get the app
Policy

OpenAI's Astra aced ExploitBench, then found two zero-days

Astra is OpenAI's first model rated Critical for cyber — 100% on ExploitBench, and it chained real zero-days on its own.

OpenAI's Astra aced ExploitBench, then found two zero-days

OpenAI's next model, Astra, is the first it has ever classified as Critical for cybersecurity under its Preparedness Framework — the tier that means the capability is dangerous enough to gate before shipping.

The numbers behind that call: one Astra configuration scored 100% on ExploitBench, the public benchmark for writing exploits against known bugs. So OpenAI built a harder internal port — 20 high-severity V8 vulnerabilities disclosed between June and August 2026, too recent to have been memorized — where Astra hit roughly a 39% arbitrary code-execution rate on about 76k output tokens, well ahead of GPT‑5.6 Sol on both capability and token efficiency. In expert-led testing against a hardened browser and OS, it found two previously unknown zero-days and turned them into a working chain that escaped the sandbox and ran commands on the host from an opened HTML file. Both are being disclosed to maintainers.

So Astra arrives "soon," but the offensive muscle doesn't arrive with it. Advanced cyber features go to a small tester group first, with defensive access widening later through Daybreak Blue. High-risk accounts get restricted responses, and chain-of-thought monitoring watches for misuse — OpenAI paused internal Astra work in August specifically to bolt on stricter safeguards.

Why it matters: the gap between "AI can spot a bug" and "AI writes the whole exploit chain unattended" just closed, and vendor policy is currently the only thing standing in it.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play