Get the app
LLMs

DeepSeek drops V4-Flash beta — beats its own Pro on agents

DeepSeek-V4-Flash-0731 is live in API beta, outscoring V4-Pro-Preview on agentic benchmarks and adding Responses API support.

DeepSeek drops V4-Flash beta — beats its own Pro on agents

DeepSeek-V4-Flash-0731 landed in API beta, and the headline claim is the odd one: the cheap, fast tier reportedly beats V4-Pro-Preview on agentic benchmarks. Flash models usually trade capability for latency. Here the distillation apparently kept the tool-calling and multi-step reasoning intact — which is exactly the part agent builders care about.

The other quiet upgrade: Flash now speaks the Responses API format, the stateful, tool-native successor to chat completions. That makes it a near drop-in for anyone whose agent scaffolding already targets that shape — swap the base URL, keep the code. Combined with DeepSeek's usual pricing posture, it's a direct shot at the cheap-agentic-workhorse slot.

V4-Pro is reportedly next, and soon. The 0731 date stamp suggests a fast-moving checkpoint cadence rather than a single frozen release, so expect the benchmark numbers to move.

Worth noting: these claims come from DeepSeek's own announcement channels, and independent evals on V4-Flash aren't in yet. Beta APIs also tend to change quotas and behavior without much warning — build accordingly.

Why it matters: if a Flash-tier model really out-agents a Pro-tier one, the cost floor for production agents just dropped again.

Sources

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play