LLM API 요금 계산기 2026
OpenAI·Anthropic·Google·xAI·DeepSeek·Mistral·Meta의 28개 모델 토큰 요금을 비교하고, 프롬프트 캐싱·배치 할인까지 반영한 실제 월 비용을 계산해 보세요.
1 · 워크로드를 입력하세요
2 · 모델별 월 비용 비교
| 벤더 | 모델 | 컨텍스트 | 입력 $/1M | 출력 $/1M | 요청당 $ | 월 비용 ▲ |
|---|
캐시 적중 입력은 각 벤더의 캐시 요금(입력의 약 10%)으로 계산되며, 캐시 요금이 공개되지 않은 모델은 슬라이더의 영향을 받지 않습니다. 배치 −50%는 지원 모델의 입력·출력 모두에 적용됩니다. Claude Sonnet 5는 2026년 8월 31일까지 프로모션 요금(정상가 $3/$15)입니다.
토큰 카운터 — 텍스트를 붙여넣으면 토큰 수를 추정합니다
영어 기준 1토큰 ≈ 4자 ≈ 0.75단어입니다. 한국어·일본어·코드는 글자당 토큰을 더 많이 사용합니다. 정확한 수치는 각 모델의 토크나이저에 따라 다릅니다.
How LLM API pricing works in 2026
Every major AI provider bills by the token — roughly a word fragment. You pay separately for input tokens (your prompt, system message, retrieved context) and output tokens (the model's reply). Output tokens cost 2–6× more than input tokens, so long-form generation is what really drives your bill.
Two discounts matter more than the sticker price. Prompt caching bills repeated prompt prefixes (system prompts, few-shot examples, long documents) at about 10% of the normal input rate — a chatbot with a 3,000-token system prompt saves dramatically at scale. The Batch API gives roughly 50% off both directions when you can wait minutes-to-hours for results, which suits evals, ETL and bulk summarization.
The spread between models is enormous: as of August 2026, a million output tokens runs from about $0.05 (Llama 3.1 8B via providers) to $180 (GPT‑5.5 Pro) — a 3,600× range. Picking the smallest model that clears your quality bar is the single biggest cost lever, ahead of caching, batching, or prompt trimming.
Frequently asked questions
How much does the GPT‑5.6 API cost?
How much does the Claude API cost?
What's the cheapest LLM API right now?
How many tokens is 1,000 words?
Do caching and batch discounts stack?
Where do these prices come from?
Go deeper
Every Anthropic model's current rates. OpenAI API pricing
GPT-5.6 Sol, Terra, Luna and the full lineup. Gemini API pricing
Google's tiers, incl. the long-context surcharge. Claude vs GPT
Tier-by-tier cost comparison with caching factored in. Cheapest LLM APIs
The 2026 budget leaderboard — and when cheap is enough. Prompt caching guide
The 90% discount most teams leave on the table.
Sponsor TokenTally
Building an AI product, inference platform, or dev tool? Put it in front of developers actively comparing LLM costs. One tasteful placement, no tracking scripts. Get in touch →