LLM Token Cost Calculator

Know exactly what your AI will cost before you build it. Compare real-time per-1M-token pricing across GPT, Claude, Gemini, and more - then estimate your real monthly spend on input and output tokens in seconds.

Dial in your token volume

1,000,000
1,000,000

Live Cost Comparison

Anthropic LogoAnthropic

Claude Opus 4.1Flagship
$90.00
/ mo
Claude Opus 4Flagship
$90.00
/ mo
Claude Fable 5.1Flagship
$60.00
/ mo
Claude Fable 5Flagship
$60.00
/ mo

DeepSeek LogoDeepSeek

DeepSeek V4 Pro 0423Mid
$2.87
/ mo
R1 Distill Llama 70BMid
$1.60
/ mo
R1Mid
$3.20
/ mo
DeepSeek V4 Pro 0813Mid
$2.32
/ mo

Google LogoGoogle

Nano Banana Pro (Gemini 3 Pro Image)Mid
$14.00
/ mo
Gemini 3.5 FlashMid
$10.50
/ mo
Gemini 2.5 ProMid
$11.25
/ mo
Gemini 3.8 FlashMid
$4.50
/ mo

OpenAI LogoOpenAI

o1-proFlagship
$750.00
/ mo
GPT-5.5 ProFlagship
$210.00
/ mo
GPT-5.4 ProFlagship
$210.00
/ mo
GPT-4Flagship
$90.00
/ mo

Qwen LogoQwen

Qwen3.8 Max (0902)Mid
$8.00
/ mo
Qwen3.8 2.4T A95BMid
$8.00
/ mo
Qwen3.7 MaxMid
$5.90
/ mo
Qwen3 Max ThinkingMid
$4.68
/ mo

xAI LogoSpaceXAI

Grok 4.6Mid
$8.00
/ mo
Grok 4.5Mid
$8.00
/ mo
Grok 4.3Mid
$3.75
/ mo
Grok 4.20 Multi-AgentMid
$3.75
/ mo

Full pricing comparison (per 1M tokens)

Provider Model Tier Input $/1M Output $/1M Blended* $/1M
AnthropicClaude Opus 4.1Flagship$15.00$75.00$30.00
AnthropicClaude Opus 4Flagship$15.00$75.00$30.00
AnthropicClaude Fable 5.1Flagship$10.00$50.00$20.00
AnthropicClaude Fable 5Flagship$10.00$50.00$20.00
AnthropicClaude Opus 5Flagship$5.00$25.00$10.00
AnthropicClaude Opus 4.8Flagship$5.00$25.00$10.00
AnthropicClaude Opus 4.7Flagship$5.00$25.00$10.00
AnthropicClaude Opus 4.6Flagship$5.00$25.00$10.00
AnthropicClaude Opus 4.5Flagship$5.00$25.00$10.00
AnthropicClaude Sonnet 4.6Flagship$3.00$15.00$6.00
AnthropicClaude Sonnet 4.5Flagship$3.00$15.00$6.00
AnthropicClaude Sonnet 4Flagship$3.00$15.00$6.00
AnthropicClaude Sonnet 5Mid$2.00$10.00$4.00
AnthropicClaude Haiku 4.5Mid$1.00$5.00$2.00
AnthropicClaude 3 HaikuBudget$0.25$1.25$0.50
DeepSeekDeepSeek V4 Pro 0423Mid$0.96$1.91$1.19
DeepSeekR1 Distill Llama 70BMid$0.80$0.80$0.80
DeepSeekR1Mid$0.70$2.50$1.15
DeepSeekDeepSeek V4 Pro 0813Mid$0.58$1.74$0.87
DeepSeekR1 0528Budget$0.50$2.15$0.91
DeepSeekDeepSeek V3Budget$0.32$0.89$0.46
DeepSeekDeepSeek V3 0324Budget$0.29$1.14$0.50
DeepSeekDeepSeek V3.2 ExpBudget$0.27$0.41$0.30
DeepSeekDeepSeek V3.1 TerminusBudget$0.27$1.00$0.45
DeepSeekDeepSeek V3.2Budget$0.27$0.40$0.30
DeepSeekDeepSeek V3.1Budget$0.25$0.95$0.42
DeepSeekDeepSeek V4 Flash 0423Budget$0.09$0.18$0.11
DeepSeekDeepSeek V4 Flash 0731Budget$0.07$0.18$0.09
GoogleNano Banana Pro (Gemini 3 Pro Image)Mid$2.00$12.00$4.50
GoogleGemini 3.5 FlashMid$1.50$9.00$3.38
GoogleGemini 2.5 ProMid$1.25$10.00$3.44
GoogleGemini 3.8 FlashMid$0.75$3.75$1.50
GoogleGemini 3.7 FlashMid$0.75$3.75$1.50
GoogleGemini 3.6 FlashMid$0.75$3.75$1.50
GoogleGemma 2 27BMid$0.65$0.65$0.65
GoogleNano Banana 2 (Gemini 3.1 Flash Image)Budget$0.50$3.00$1.13
GoogleGemini 3.5 Flash LiteBudget$0.30$2.50$0.85
GoogleNano Banana (Gemini 2.5 Flash Image)Budget$0.30$2.50$0.85
GoogleGemini 2.5 FlashBudget$0.30$2.50$0.85
GoogleNano Banana 2 Lite (Gemini 3.1 Flash Lite Image)Budget$0.25$1.50$0.56
GoogleGemini 3.1 Flash LiteBudget$0.25$1.50$0.56
GoogleGemini 2.5 Flash LiteBudget$0.10$0.40$0.17
GoogleGemma 4 31BBudget$0.09$0.34$0.15
GoogleGemma 3 27BBudget$0.08$0.45$0.17
GoogleGemma 4 26B A4B Budget$0.07$0.34$0.14
GoogleGemma 3 4BBudget$0.05$0.10$0.06
GoogleGemma 3 12BBudget$0.05$0.15$0.07
OpenAIo1-proFlagship$150.00$600.00$262.50
OpenAIGPT-5.5 ProFlagship$30.00$180.00$67.50
OpenAIGPT-5.4 ProFlagship$30.00$180.00$67.50
OpenAIGPT-4Flagship$30.00$60.00$37.50
OpenAIGPT-5.2 ProFlagship$21.00$168.00$57.75
OpenAIo3 ProFlagship$20.00$80.00$35.00
OpenAIGPT-5 ProFlagship$15.00$120.00$41.25
OpenAIo1Flagship$15.00$60.00$26.25
OpenAIGPT-6 AstraFlagship$10.00$50.00$20.00
OpenAIGPT-6 Astra ProFlagship$10.00$50.00$20.00
OpenAIGPT-5 ImageFlagship$10.00$10.00$10.00
OpenAIGPT-4 TurboFlagship$10.00$30.00$15.00
OpenAIGPT-5.4 Image 2Flagship$8.00$15.00$9.75
OpenAIGPT-5.5Flagship$5.00$30.00$11.25
OpenAIGPT-4o (2024-05-13)Flagship$5.00$15.00$7.50
OpenAIGPT-3.5 Turbo 16kFlagship$3.00$4.00$3.25
OpenAIGPT-5.4Mid$2.50$15.00$5.63
OpenAIGPT AudioMid$2.50$10.00$4.38
OpenAIGPT-5 Image MiniMid$2.50$2.00$2.38
OpenAIGPT-4o (2024-11-20)Mid$2.50$10.00$4.38
OpenAIGPT-4o (2024-08-06)Mid$2.50$10.00$4.38
OpenAIGPT-4oMid$2.50$10.00$4.38
OpenAIGPT-5.6 Terra ProMid$2.00$12.00$4.50
OpenAIGPT-5.6 TerraMid$2.00$12.00$4.50
OpenAIGPT-5.6 Sol ProMid$2.00$10.00$4.00
OpenAIGPT-5.6 SolMid$2.00$10.00$4.00
OpenAIo3Mid$2.00$8.00$3.50
OpenAIGPT-4.1Mid$2.00$8.00$3.50
OpenAIGPT-5.3-CodexMid$1.75$14.00$4.81
OpenAIGPT-5.2-CodexMid$1.75$14.00$4.81
OpenAIGPT-5.2 ChatMid$1.75$14.00$4.81
OpenAIGPT-5.2Mid$1.75$14.00$4.81
OpenAIGPT-5.1-Codex-MaxMid$1.25$10.00$3.44
OpenAIGPT-5.1Mid$1.25$10.00$3.44
OpenAIGPT-5.1-CodexMid$1.25$10.00$3.44
OpenAIGPT-5Mid$1.25$10.00$3.44
OpenAIo4 Mini HighMid$1.10$4.40$1.93
OpenAIo4 MiniMid$1.10$4.40$1.93
OpenAIo3 Mini HighMid$1.10$4.40$1.93
OpenAIo3 MiniMid$1.10$4.40$1.93
OpenAIGPT-3.5 Turbo (older v0613)Mid$1.00$2.00$1.25
OpenAIGPT-5.4 MiniMid$0.75$4.50$1.69
OpenAIGPT Audio MiniMid$0.60$2.40$1.05
OpenAIGPT-3.5 TurboBudget$0.50$1.50$0.75
OpenAIGPT-4.1 MiniBudget$0.40$1.60$0.70
OpenAIGPT-5.1-Codex-MiniBudget$0.25$2.00$0.69
OpenAIGPT-5 MiniBudget$0.25$2.00$0.69
OpenAIGPT-5.6 Luna ProBudget$0.20$1.20$0.45
OpenAIGPT-5.6 LunaBudget$0.20$1.20$0.45
OpenAIGPT-5.4 NanoBudget$0.20$1.25$0.46
OpenAIGPT-4o-miniBudget$0.15$0.60$0.26
OpenAIGPT-4o-mini (2024-07-18)Budget$0.15$0.60$0.26
OpenAIGPT-4.1 NanoBudget$0.10$0.40$0.17
OpenAIgpt-oss-safeguard-20bBudget$0.07$0.30$0.13
OpenAIGPT-5 NanoBudget$0.05$0.40$0.14
OpenAIgpt-oss-120bBudget$0.04$0.17$0.07
OpenAIgpt-oss-20bBudget$0.03$0.13$0.06
QwenQwen3.8 Max (0902)Mid$2.00$6.00$3.00
QwenQwen3.8 2.4T A95BMid$2.00$6.00$3.00
QwenQwen3.7 MaxMid$1.48$4.42$2.21
QwenQwen3 Max ThinkingMid$0.78$3.90$1.56
QwenQwen3 MaxMid$0.78$3.90$1.56
QwenQwen3 Coder PlusMid$0.65$3.25$1.30
QwenQwen3.5 397B A17BMid$0.55$3.50$1.29
QwenQwen3 235B A22BBudget$0.45$1.82$0.80
QwenQwen3.8 27BBudget$0.42$3.00$1.06
QwenQwen3 VL 235B A22B ThinkingBudget$0.40$4.00$1.30
QwenQwen3.6 PlusBudget$0.33$1.95$0.73
QwenQwen3.7 PlusBudget$0.32$1.28$0.56
QwenQwen3.5-35B-A3BBudget$0.31$1.25$0.55
QwenQwen3.5 Plus 2026-04-20Budget$0.30$1.80$0.67
QwenQwen3.6 27BBudget$0.30$2.00$0.72
QwenQwen3 Coder 480B A35BBudget$0.30$1.00$0.47
QwenQwen3.5-122B-A10BBudget$0.29$2.40$0.82
QwenQwen3.5 Plus 2026-02-15Budget$0.26$1.56$0.58
QwenQwen Plus 0728Budget$0.26$0.78$0.39
QwenQwen-PlusBudget$0.26$0.78$0.39
QwenQwen3 235B A22B Thinking 2507Budget$0.23$2.30$0.75
QwenQwen3 14BBudget$0.23$0.91$0.40
QwenQwen3 VL 30B A3B ThinkingBudget$0.20$2.40$0.75
QwenQwen3 30B A3B Thinking 2507Budget$0.20$2.40$0.75
QwenQwen3.5-27BBudget$0.20$1.56$0.54
QwenQwen3 Coder FlashBudget$0.20$0.97$0.39
QwenQwen3.6 FlashBudget$0.19$1.13$0.42
QwenQwen3 VL 8B ThinkingBudget$0.18$2.10$0.66
QwenQwen3.8 FlashBudget$0.15$0.47$0.23
QwenQwen3 Next 80B A3B ThinkingBudget$0.15$1.20$0.41
QwenQwen3 Coder NextBudget$0.12$0.80$0.29
QwenQwen3 30B A3BBudget$0.12$0.50$0.21
QwenQwen3 8BBudget$0.12$0.45$0.20
QwenQwen3.6 35B A3BBudget$0.10$0.90$0.30
QwenQwen3.5-9BBudget$0.10$0.15$0.11
QwenQwen3 32BBudget$0.08$0.28$0.13
QwenQwen3.5-FlashBudget$0.07$0.26$0.11
QwenQwen3.7 FlashBudget$0.03$0.13$0.06
SpaceXAIGrok 4.6Mid$2.00$6.00$3.00
SpaceXAIGrok 4.5Mid$2.00$6.00$3.00
SpaceXAIGrok 4.3Mid$1.25$2.50$1.56
SpaceXAIGrok 4.20 Multi-AgentMid$1.25$2.50$1.56
SpaceXAIGrok 4.20Mid$1.25$2.50$1.56
SpaceXAIGrok Build 0.1Mid$1.00$2.00$1.25

*Blended = 3 input : 1 output, a common chat ratio. Prices are dynamic and fetched live. Cache misses are shown; cached input is typically ~10% of base rate on supported providers.

Pricing Data Source

Our calculator now fetches live pricing data directly from the OpenRouter API. This means the prices above are always up-to-date with the latest market rates for both input and output tokens.


What Is a Token, Anyway?

A token is the smallest unit of text a model reads or writes - roughly ¾ of a word. "The quick brown fox" ≈ 4 tokens. Pricing is quoted per 1 million tokens because real workloads (chat apps, agents, document processing) routinely consume millions of tokens a month.

Two things drive your bill:

  1. Input tokens - everything you send: your prompt, conversation history, retrieved context (RAG), system instructions.
  2. Output tokens - everything the model generates, including any hidden "thinking" tokens in reasoning models.

The rule of thumb: output is always more expensive than input - anywhere from 2× to 5×. A chatbot that generates long answers will burn far more budget on output than input, even though the prompt feels like the "big" part.


How to Cut Your Token Costs (Without Cutting Quality)

  1. Cache your context. Most providers price cached input at ~10% of the base rate. If the same system prompt + documents are sent repeatedly, caching can cut input cost by 90%.
  2. Batch non-urgent jobs. OpenAI and xAI offer ~50% off for async/batch processing. Perfect for summarization, classification, and backfill jobs.
  3. Right-size the model. Reserve flagships for reasoning-heavy work. Route routine tasks to faster, budget-friendly alternatives.
  4. Watch output tokens. Reasoning models burn hidden "thinking" tokens on output. Cap max_tokens and tune prompts to answer concisely.
  5. Go private for high volume. Past a few hundred million tokens a month, self-hosted open models (Qwen, DeepSeek, Llama, Mistral) on your own GPU typically undercut API pricing - Vistaran can run that for you.

FAQ

Why is output more expensive than input?
Output tokens are generated one at a time and can't be cached or parallelized the way input can. Reasoning models add hidden "thinking" tokens on the output side, further inflating the ratio.
What are cached input tokens?
When the same prompt prefix is sent repeatedly, the provider stores and reuses the computed representation, charging a fraction (~10%) of the normal input rate. This is the single biggest lever for RAG and chatbot cost.
Do these prices include a "thinking" / reasoning surcharge?
For most providers, thinking tokens are billed as output tokens at the standard output rate (Gemini and Grok bill them this way). Always check whether a reasoning model's quoted price includes or excludes thinking.
Is this calculator for API pricing or ChatGPT/Claude Pro subscriptions?
API pricing - the per-token rates you pay when building your own product on these models. Consumer subscriptions (ChatGPT Plus, Claude Pro) are flat monthly fees and aren't comparable per token.
LET'S CONNECT

Build on the Right Model - and the Right Stack

Picking a model is only step one. The bigger decision is whether you run it on a public API, on dedicated infrastructure, or on-premise for compliance. That's exactly what we engineer every day.

AI AGENT AS A SERVICEPRIVATE LLM HOSTINGAI AUTOMATION AUDIT
BUILT BY EXPERTS
Engineers shipping AI since 2017
Boomerr.ai, Scanview, and Compliance View in production
Secure, private deployments