Get the app
Ethics

Anthropic researcher quits: labs are 'gambling with our lives'

Jacob Coxon quit Anthropic saying labs are racing to superintelligence. Its alignment lead replied: >10% odds AI kills all humans within a decade.

Anthropic researcher quits: labs are 'gambling with our lives'

Anthropic's own alignment lead, Evan Hubinger, answered a colleague's resignation with a number rather than a rebuttal. He wrote: "we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade."

The person who resigned is Jacob Coxon, who spent three years doing pretraining research at OpenAI and then Anthropic. His exit post: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." He says the people building frontier models "earnestly believe that it could kill us all by the end of the decade." According to him, they say this in private while sounding measured in public. He is calling for a temporary ban on improving model capabilities.

Hubinger didn't fully side with him. He said Anthropic is "trying its best", but admitted the lab does "not yet have a plan to solve alignment for superintelligence" and is "not clearly on track to." Neither Anthropic nor OpenAI has put out an official response.

Why it matters: when a frontier lab's alignment lead puts extinction odds in double digits in public, it gets much harder to wave the risk off as "doomer" talk.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play