Get the app
Ethics

Anthropic safety researchers are quitting — and saying why

Three AI safety researchers walked out of Anthropic and Google DeepMind in a week, warning the labs are racing past their own guardrails.

Jacob Coxon quit Anthropic on 9 September and posted his reasons publicly: the labs are "racing straight to self-improving superintelligence and gambling with our lives." He had spent roughly three years on model-training research across OpenAI and Anthropic, and his verdict on both was blunt — "neither company is acting responsibly." The post cleared 150 million views.

Then it kept going. Joe Benton, who led a safety research team at Anthropic, left days later, warning that AI progress could go "from merely blistering at the minute to uncontrollable" and noting that "basically all of the transparency about these risks that is coming from the companies is entirely voluntary." Josh Engels walked out of Google DeepMind the same week with a sharper line: "There are no adults in the room. People are trying their best, but there is no one coming to save us."

The awkward part for Anthropic is that safety is the brand. Two current employees — alignment science lead Evan Hubinger and Cognitive Oversight lead Samuel Marks — publicly backed pieces of Coxon's argument instead of rebutting them, with Marks noting he posted in a personal capacity. The company's only official line so far: "We have always been transparent that AI will bring both enormous benefits and unprecedented risks."

Why it matters: voluntary self-regulation is the entire governance model for frontier AI right now, and the people inside the labs are the ones saying it isn't holding.

Sources

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play