Claude API pricing in 2026: every model, current rates
Anthropic prices the Claude API per million tokens, with separate input and output rates. As of August 2026 the lineup spans four tiers, from Haiku for high-volume light work to Fable 5, the frontier model. All tiers support prompt caching (~90% off repeated input) and a 50% Batch API discount.
| Model | Input | Cached input | Output | Context |
|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $1.00 | $50.00 | 1M |
| Claude Opus 5 | $5.00 | $0.50 | $25.00 | 1M |
| Claude Sonnet 5 * | $2.00 | $0.20 | $10.00 | 1M |
| Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | 200K |
* Sonnet 5 promotional rate through August 31, 2026; standard rate $3 input / $15 output. Cache hits bill at 10% of input; Batch API takes 50% off both directions.
Which Claude model should you pay for?
Sonnet 5 is the default for production apps — near-flagship quality at 1/5th the frontier price, and the promo rate makes August an especially good month to lock in usage. Haiku 4.5 handles classification, extraction and routing at a fifth of Sonnet's price. Opus 5 and Fable 5 earn their premium on hard reasoning and agentic coding, where fewer retries can beat a cheaper per-token rate — measure per-task cost, not per-token.
Three ways to cut your Claude bill
First, prompt caching: mark your stable system prompt with cache_control and repeated input bills at 10%. A chatbot with a 3K-token system prompt saves ~50% overall. Second, the Batch API: 50% off everything if you can wait for async results. Third, right-size the model — run an eval on Haiku before assuming you need Sonnet. Details with worked math in our caching guide and Claude vs GPT comparison.
→ What would YOUR workload cost?Enter your token mix once — compare all 28 models side by side, with caching and batch discounts.