Get the app
LLMs

Three labs, one week: AI pricing just split in two directions

Google halved Gemini 3.7 Flash to $0.75/M, OpenAI shipped a 14x-faster tier, and DeepSeek hiked prices 4.5x.

Three labs, one week: AI pricing just split in two directions

Gemini 3.7 Flash went GA on Aug 13 at $0.75/M input and $3.75/M output — half of 3.6 Flash. The catch: that's introductory pricing through Dec 31, 2026. On Jan 1 it reverts to $1.50/$7.50. Google is pitching it as the workhorse for coding and agents, and claims it tops Claude on business workflows.

OpenAI's Ultrafast goes the other way — not cheaper, just faster. GPT-5.6 Sol on Cerebras wafer-scale chips hits 750 output tokens/sec, up to 14x standard mode, with no quality drop (44GB of on-chip SRAM sidesteps the GPU memory-bandwidth wall). Limited preview. No price, no date.

DeepSeek V4-Pro left preview the same week — 1M-token context, 384K max output, agent-tuned, with its agent software open-sourced. Then came the bill: from 16:00 UTC on Aug 16, peak output goes from $0.87 to $3.96/M, roughly 4.5x, plus new peak/off-peak billing (off-peak is half price). The industry's most reliable price floor just moved up.

Why it matters: "AI gets cheaper every quarter" is no longer a safe assumption — cost, speed, and context are now three separate things you pay for separately.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play