DeepSeek V4.1 Flash: price, context window and capabilities
DeepSeek V4.1 Flash is a model from DeepSeek (deepseek/deepseek-v4.1-flash). It supports reasoning. It supports tool use. It costs $0.003 per million input tokens and $2.40 per million output tokens, with a 1.0M-token context window.
What it costs in practice
Using the listed prices: a chat turn of 2,000 input and 500 output tokens costs about $0.0012, so 10,000 such turns cost about $12.06. Summarising a 100,000-token document into 2,000 tokens costs about $0.0051. A long agent run that reads 1 million tokens and writes 50,000 costs about $0.123.
How it compares
Among the 40 models we track, DeepSeek V4.1 Flash ranks #1 by input price and #18 by output price (1 is cheapest), and #6 by context window (1 is largest). Models with the same tool and reasoning support that cost less per output token: DeepSeek V4 Pro 0423 ($0.418), Qwen3.8 Omni Flash ($0.470), Qwen3.8 Flash ($0.470), GPT-6 Luna Pro ($0.500). Models with a larger context window at the same or lower output price: Llama 4 Scout (1.3M), GPT-6 Luna Pro (1.1M), GPT-6 Luna (1.1M).
More from DeepSeek
- DeepSeek V4 Pro 0813: $0.170 in / $4.20 out, 1.0M context
- DeepSeek V4 Flash 0731: $0.015 in / $1.28 out, 1.0M context
- DeepSeek V4 Pro 0423: $0.209 in / $0.418 out, 1.0M context
Prices and limits come from OpenRouter's public model catalogue, last synced 4 Oct 2026. Provider pricing can differ from OpenRouter's and changes often; confirm on the provider's own page before you commit. See the guide to choosing a model or the full comparison.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.