Get the app
LLMs

Three labs, one week: AI pricing just split in two directions

Google halved Gemini 3.7 Flash to $0.75/M, OpenAI shipped a 14x-faster tier, and DeepSeek hiked prices 4.5x.

Three labs, one week: AI pricing just split in two directions

Gemini 3.7 Flash went GA on Aug 13 at $0.75/M input and $3.75/M output — half of 3.6 Flash. The catch: that's introductory pricing through Dec 31, 2026. On Jan 1 it reverts to $1.50/$7.50. Google is pitching it as the workhorse for coding and agents, and claims it tops Claude on business workflows.

OpenAI's Ultrafast goes the other way — not cheaper, just faster. GPT-5.6 Sol on Cerebras wafer-scale chips hits 750 output tokens/sec, up to 14x standard mode, with no quality drop (44GB of on-chip SRAM sidesteps the GPU memory-bandwidth wall). Limited preview. No price, no date.

DeepSeek V4-Pro left preview the same week — 1M-token context, 384K max output, agent-tuned, with its agent software open-sourced. Then came the bill: from 16:00 UTC on Aug 16, peak output goes from $0.87 to $3.96/M, roughly 4.5x, plus new peak/off-peak billing (off-peak is half price). The industry's most reliable price floor just moved up.

Why it matters: "AI gets cheaper every quarter" is no longer a safe assumption — cost, speed, and context are now three separate things you pay for separately.

Sources

Primary: the company, paper or repository

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play