Anthropic safety researchers are quitting — and saying why
Three AI safety researchers walked out of Anthropic and Google DeepMind in a week, warning the labs are racing past their own guardrails.
Jacob Coxon quit Anthropic on 9 September and posted his reasons publicly: the labs are "racing straight to self-improving superintelligence and gambling with our lives." He had spent roughly three years on model-training research across OpenAI and Anthropic, and his verdict on both was blunt — "neither company is acting responsibly." The post cleared 150 million views.
Then it kept going. Joe Benton, who led a safety research team at Anthropic, left days later, warning that AI progress could go "from merely blistering at the minute to uncontrollable" and noting that "basically all of the transparency about these risks that is coming from the companies is entirely voluntary." Josh Engels walked out of Google DeepMind the same week with a sharper line: "There are no adults in the room. People are trying their best, but there is no one coming to save us."
The awkward part for Anthropic is that safety is the brand. Two current employees — alignment science lead Evan Hubinger and Cognitive Oversight lead Samuel Marks — publicly backed pieces of Coxon's argument instead of rebutting them, with Marks noting he posted in a personal capacity. The company's only official line so far: "We have always been transparent that AI will bring both enormous benefits and unprecedented risks."
Why it matters: voluntary self-regulation is the entire governance model for frontier AI right now, and the people inside the labs are the ones saying it isn't holding.
Sources
Independent coverage
- Anthropic researcher resigns, warning that AI companies are 'gambling with our lives' fortune.com
- Two AI researchers leave Anthropic, Google over safety concerns nbcnews.com
- Scoop: Anthropic whistleblower gave up his equity to leave the company axios.com
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.