Which LLM should I use?
The right model depends on the job, so these lists answer the common questions from current API prices: what is cheapest, what reads the most, and what is the best value when you need reasoning and tools. They are computed from the 40 models we track and refresh weekly (last synced 4 Oct 2026).
The cheapest models that support tool use
Tool use (function calling) is what lets a model drive an agent or call your APIs. Ranked by price per million output tokens, since output is usually the larger cost.
| Model | Input | Output | Context |
|---|---|---|---|
| Ministral 3 14B 2512 Mistral AI | $0.200 | $0.200 | 262K |
| Llama 4 Scout Meta | $0.100 | $0.300 | 1.3M |
| DeepSeek V4 Pro 0423 DeepSeek | $0.209 | $0.418 | 1.0M |
| Qwen3.8 Omni Flash Qwen (Alibaba) | $0.150 | $0.470 | 1M |
| Qwen3.8 Flash Qwen (Alibaba) | $0.150 | $0.470 | 1M |
| GPT-6 Luna Pro OpenAI | $0.100 | $0.500 | 1.1M |
The cheapest reasoning models
Reasoning models spend extra tokens thinking before they answer, which helps on maths, code and multi-step tasks. These support both reasoning and tools.
| Model | Input | Output | Context |
|---|---|---|---|
| DeepSeek V4 Pro 0423 DeepSeek | $0.209 | $0.418 | 1.0M |
| Qwen3.8 Omni Flash Qwen (Alibaba) | $0.150 | $0.470 | 1M |
| Qwen3.8 Flash Qwen (Alibaba) | $0.150 | $0.470 | 1M |
| GPT-6 Luna Pro OpenAI | $0.100 | $0.500 | 1.1M |
| GPT-6 Luna OpenAI | $0.100 | $0.500 | 1.1M |
| GLM 5.3 Flash Z.ai | $0.150 | $0.500 | 1.0M |
The largest context windows
Context is how much text a model can consider at once: whole codebases, long contracts, large document sets. A bigger window costs more to fill, so check the input price too.
| Model | Input | Output | Context |
|---|---|---|---|
| Llama 4 Scout Meta | $0.100 | $0.300 | 1.3M |
| GPT-6.1 Sol Pro OpenAI | $2.00 | $10.00 | 1.1M |
| GPT-6.1 Sol OpenAI | $2.00 | $10.00 | 1.1M |
| GPT-6 Luna Pro OpenAI | $0.100 | $0.500 | 1.1M |
| GPT-6 Luna OpenAI | $0.100 | $0.500 | 1.1M |
| Gemini 3.8 Flash Google | $0.750 | $3.75 | 1.0M |
Best value for serious work
Models with reasoning, tool use and at least a 200K-token context, cheapest first by output price. A reasonable shortlist to test on your own task.
| Model | Input | Output | Context |
|---|---|---|---|
| DeepSeek V4 Pro 0423 DeepSeek | $0.209 | $0.418 | 1.0M |
| Qwen3.8 Omni Flash Qwen (Alibaba) | $0.150 | $0.470 | 1M |
| Qwen3.8 Flash Qwen (Alibaba) | $0.150 | $0.470 | 1M |
| GPT-6 Luna Pro OpenAI | $0.100 | $0.500 | 1.1M |
| GPT-6 Luna OpenAI | $0.100 | $0.500 | 1.1M |
| GLM 5.3 Flash Z.ai | $0.150 | $0.500 | 1.0M |
By provider
- Anthropic: 4 models, from $2.00 to $50.00 per million tokens. Claude Sonnet 5.5, Claude Opus 5.5, Claude Fable 5.1
- OpenAI: 4 models, from $0.100 to $10.00 per million tokens. GPT-6.1 Sol Pro, GPT-6.1 Sol, GPT-6 Luna Pro
- Google: 4 models, from $0.300 to $3.75 per million tokens. Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash
- Meta: 4 models, from $0.027 to $0.652 per million tokens. Llama 4 Maverick, Llama 4 Scout, Llama 3.3 70B Instruct
- DeepSeek: 4 models, from $0.003 to $4.20 per million tokens. DeepSeek V4.1 Flash, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731
- Mistral AI: 4 models, from $0.150 to $7.50 per million tokens. Mistral Medium 3.5, Mistral Small 4, Devstral 2 2512
- Qwen (Alibaba): 4 models, from $0.150 to $12.00 per million tokens. Qwen3.8 Max Prime, Qwen3.8 Omni Flash, Qwen3.8 Max (0902)
- xAI: 4 models, from $1.00 to $6.00 per million tokens. Grok 4.7, Grok 4.6, Grok 4.5
- Moonshot AI: 4 models, from $0.450 to $13.00 per million tokens. Kimi K3, Kimi K2.7 Code, Kimi K2.6
- Z.ai: 4 models, from $0.150 to $8.80 per million tokens. GLM 5.3 Prime, GLM 5.3 FlashX, GLM 5.3 Flash
How to choose
Price is only one input. Run your own prompts on two or three candidates and compare quality, latency and cost on your task, because benchmark rankings rarely predict that for you. Prices here are OpenRouter's listings and can differ from a provider's direct pricing; see the full comparison table for every model.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.