Get the app
LLMs

DeepSeek drops V4-Flash beta — beats its own Pro on agents

DeepSeek-V4-Flash-0731 is live in API beta, outscoring V4-Pro-Preview on agentic benchmarks and adding Responses API support.

DeepSeek drops V4-Flash beta — beats its own Pro on agents

DeepSeek-V4-Flash-0731 landed in API beta, and the headline claim is the odd one: the cheap, fast tier reportedly beats V4-Pro-Preview on agentic benchmarks. Flash models usually trade capability for latency. Here the distillation apparently kept the tool-calling and multi-step reasoning intact — which is exactly the part agent builders care about.

The other quiet upgrade: Flash now speaks the Responses API format, the stateful, tool-native successor to chat completions. That makes it a near drop-in for anyone whose agent scaffolding already targets that shape — swap the base URL, keep the code. Combined with DeepSeek's usual pricing posture, it's a direct shot at the cheap-agentic-workhorse slot.

V4-Pro is reportedly next, and soon. The 0731 date stamp suggests a fast-moving checkpoint cadence rather than a single frozen release, so expect the benchmark numbers to move.

Worth noting: these claims come from DeepSeek's own announcement channels, and independent evals on V4-Flash aren't in yet. Beta APIs also tend to change quotas and behavior without much warning — build accordingly.

Why it matters: if a Flash-tier model really out-agents a Pro-tier one, the cost floor for production agents just dropped again.

Sources

Written by an AI pipeline from the sources above. How it works.

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play