Get the app
Industry

DeepSeek ends the price war: API rates jump up to 4.7x

From Aug 16, DeepSeek doubles off-peak API prices and doubles again at peak. V4-Flash output: $0.28 → $1.32 per million tokens.

DeepSeek ends the price war: API rates jump up to 4.7x

The cheap-tokens era has an expiry date: 16:00 UTC on August 16. DeepSeek is moving both V4 models to peak / off-peak billing, and off-peak is already roughly double what you pay today.

The numbers, per million tokens. V4-Flash output goes from $0.28 to $0.66 off-peak and $1.32 at peak — a 4.7x jump at the top end. Cache-miss input climbs $0.14 → $0.22 / $0.44. V4-Pro output moves $0.87 → $1.98 / $3.96, with cache-hit input hit hardest: $0.003625 → $0.044 at peak, an 1,100% increase. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; everything else is off-peak, priced at exactly half.

DeepSeek flagged a "significant" hike on August 6 without numbers, then dropped the table alongside the V4-Pro launch on the 13th. The stated reason: demand isn't evenly distributed, and time-of-use pricing pushes batch and non-urgent workloads out of the crunch. Translation — inference capacity is the constraint, and the company that detonated the price war is done subsidising it. If your pipeline is latency-tolerant, a cron shift out of those seven peak hours halves your bill for free.

Why it matters: the floor under LLM API pricing was never real economics — it was a land grab, and the land has been taken.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play