DeepSeek ends the price war: API rates jump up to 4.7x
From Aug 16, DeepSeek doubles off-peak API prices and doubles again at peak. V4-Flash output: $0.28 → $1.32 per million tokens.

The cheap-tokens era has an expiry date: 16:00 UTC on August 16. DeepSeek is moving both V4 models to peak / off-peak billing, and off-peak is already roughly double what you pay today.
The numbers, per million tokens. V4-Flash output goes from $0.28 to $0.66 off-peak and $1.32 at peak — a 4.7x jump at the top end. Cache-miss input climbs $0.14 → $0.22 / $0.44. V4-Pro output moves $0.87 → $1.98 / $3.96, with cache-hit input hit hardest: $0.003625 → $0.044 at peak, an 1,100% increase. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; everything else is off-peak, priced at exactly half.
DeepSeek flagged a "significant" hike on August 6 without numbers, then dropped the table alongside the V4-Pro launch on the 13th. The stated reason: demand isn't evenly distributed, and time-of-use pricing pushes batch and non-urgent workloads out of the crunch. Translation — inference capacity is the constraint, and the company that detonated the price war is done subsidising it. If your pipeline is latency-tolerant, a cron shift out of those seven peak hours halves your bill for free.
Why it matters: the floor under LLM API pricing was never real economics — it was a land grab, and the land has been taken.
Sources
- Models & Pricing | DeepSeek API Docs api-docs.deepseek.com
- DeepSeek raises V4 API prices by up to 1,100% just as it launches V4-Pro techstartups.com
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.