DeepSeek drops V4-Flash beta — beats its own Pro on agents
DeepSeek-V4-Flash-0731 is live in API beta, outscoring V4-Pro-Preview on agentic benchmarks and adding Responses API support.

DeepSeek-V4-Flash-0731 landed in API beta, and the headline claim is the odd one: the cheap, fast tier reportedly beats V4-Pro-Preview on agentic benchmarks. Flash models usually trade capability for latency. Here the distillation apparently kept the tool-calling and multi-step reasoning intact — which is exactly the part agent builders care about.
The other quiet upgrade: Flash now speaks the Responses API format, the stateful, tool-native successor to chat completions. That makes it a near drop-in for anyone whose agent scaffolding already targets that shape — swap the base URL, keep the code. Combined with DeepSeek's usual pricing posture, it's a direct shot at the cheap-agentic-workhorse slot.
V4-Pro is reportedly next, and soon. The 0731 date stamp suggests a fast-moving checkpoint cadence rather than a single frozen release, so expect the benchmark numbers to move.
Worth noting: these claims come from DeepSeek's own announcement channels, and independent evals on V4-Flash aren't in yet. Beta APIs also tend to change quotas and behavior without much warning — build accordingly.
Why it matters: if a Flash-tier model really out-agents a Pro-tier one, the cost floor for production agents just dropped again.
Sources
- DeepSeek API Docs — News & Model Updates api-docs.deepseek.com
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.