Get the app
Policy

Anthropic will bill you for Claude's refusals in 3 risky areas

Anthropic is bringing back charges for requests its safety classifiers block in biology, distillation and frontier-LLM work, to make probing cost money.

Anthropic will bill you for Claude's refusals in 3 risky areas

Getting refused by Claude will cost money again, at least in some areas. Anthropic says it will resume billing input tokens for requests its safeguards block before Claude writes anything. This applies only when a classifier tags the request with one of three labels: biology safety, model distillation attacks or frontier LLM development. You still pay nothing for output, because none is generated.

The reason is distillation. Anthropic's September threat report, published 10 Sep, says it identified seven PRC-based labs that ran industrial-scale campaigns to extract Claude's abilities, with more than 200 million exchanges between them. Free refusals made that kind of automated probing cheap. Billing for them means every blocked attempt now costs the attacker.

Anthropic says the classifiers are tuned to a false positive rate below 0.1%, and that 99.7% of accounts it tested hit no billable blocks at all. For legitimate researchers who land in that small fraction, a wrong block now costs money as well as time. Claude Code users can report misfires with /feedback. Anthropic hasn't given a date for when billing starts.

Why it matters: refusals used to be free to trigger, and Anthropic is now using pricing as a security tool against people trying to copy its models.

Sources

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play