DeepSeek API pricing in 2026: the budget benchmark
DeepSeek is the price floor most teams compare everything else against. Its two current models undercut Western flagship APIs by one to two orders of magnitude on list price โ and its cache-hit pricing is the most aggressive in the industry.
| Model | Input | Cached input | Output | Context |
|---|---|---|---|---|
| DeepSeek V4 Pro | $0.435 | $0.0036 | $0.87 | 128K |
| DeepSeek V4 Flash | $0.14 | $0.0028 | $0.28 | 128K |
Cache-hit input bills at under 1% of the standard input rate โ far deeper than the ~10% cached rate typical at OpenAI, Anthropic and Google.
How it compares to GPT, Claude and Gemini
| Model | Input | Output | Output vs V4 Pro |
|---|---|---|---|
| DeepSeek V4 Pro | $0.435 | $0.87 | 1ร |
| Gemini 3 Flash | $0.50 | $3.00 | 3.4ร |
| Claude Sonnet 5 | $2.00 | $10.00 | 11ร |
| GPT-5.6 Terra | $2.00 | $12.00 | 14ร |
| Claude Fable 5 | $10.00 | $50.00 | 57ร |
On raw list price the gap is enormous. The honest caveats: DeepSeek's 128K context window is smaller than the 1M windows now standard on flagship APIs, there is no separate batch tier, and for the hardest reasoning tasks the frontier models still lead on quality. For chat, classification, summarization, extraction and most agent plumbing, many teams find V4-class quality more than sufficient.
The caching trick that changes the math
DeepSeek bills repeated prompt prefixes at $0.0036 per million tokens โ effectively free. A chatbot with a long system prompt and 80% cache-hit rate pays almost nothing for its instruction overhead. If your workload is prefix-heavy, DeepSeek's effective price drops even further below the sticker price. Model your exact mix โ cache slider included โ in the calculator, or see how it ranks in the cheapest APIs list.
โ What would YOUR workload cost?Enter your token mix once โ compare all 28 models side by side, with caching and batch discounts.