LLM API Pricing Comparison — 63 Models, Updated August 2026
The cheapest model we track is Llama 3.1 8B at $0.058 per million blended tokens; the most expensive is GPT-5.4 Pro at $67.50 per million — a 1174x spread. Every price below is sourced directly from the provider's official pricing page and dated with the last verification. Use the calculator to estimate your own monthly cost across all 63 models.
| Model | Provider | $ / M input | $ / M output | Blended ↓ | Your cost / mo |
|---|---|---|---|---|---|
| Amazon Nova Micro | Amazon | $0.03 | $0.14 | $0.06 | $0.07 |
| Amazon Nova Lite | Amazon | $0.06 | $0.24 | $0.11 | $0.12 |
| GPT-OSS 20B | Groq | $0.07 | $0.30 | $0.13 | $0.15 |
| Ministral 8B | Mistral | $0.15 | $0.15 | $0.15 | $0.19 |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | $0.17 | $0.21 |
| GPT-OSS 120B | Groq | $0.15 | $0.60 | $0.26 | $0.30 |
| Mistral Small 3.1 | Mistral | $0.15 | $0.60 | $0.26 | $0.30 |
| Codestral | Mistral | $0.30 | $0.90 | $0.45 | $0.53 |
| GPT-OSS 120B (Cerebras) | Cerebras | $0.35 | $0.75 | $0.45 | $0.54 |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | $0.45 | $0.50 |
| DeepSeek V4 Pro | DeepSeek | $0.43 | $0.87 | $0.54 | $0.65 |
| Gemini 3.5 Flash Lite | $0.25 | $1.50 | $0.56 | $0.63 | |
| Mistral Medium 3 | Mistral | $0.40 | $2.00 | $0.80 | $0.90 |
| Qwen 3.7 Plus | Qwen | $0.80 | $2.00 | $1.10 | $1.30 |
| Qwen 3.8 30B | Groq | $0.60 | $3.00 | $1.20 | $1.35 |
| Amazon Nova Pro | Amazon | $0.80 | $3.20 | $1.40 | $1.60 |
| Grok 4.3 | xAI | $1.25 | $2.50 | $1.56 | $1.88 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $2.00 | $2.25 |
| GLM-5.2 | Z.ai | $1.40 | $4.40 | $2.15 | $2.50 |
| GLM 4.7 (Cerebras) | Cerebras | $2.25 | $2.75 | $2.38 | $2.94 |
| Qwen 3.8 Max | Qwen | $1.60 | $6.40 | $2.80 | $3.20 |
| Qwen 3.7 Max | Qwen | $1.60 | $6.40 | $2.80 | $3.20 |
| Grok-4.20 Reasoning | xAI | $2.00 | $6.00 | $3.00 | $3.50 |
| Grok-4.20 | xAI | $2.00 | $6.00 | $3.00 | $3.50 |
| Mistral Large 3 | Mistral | $2.00 | $6.00 | $3.00 | $3.50 |
| Gemini 3.6 Flash | $1.50 | $9.00 | $3.38 | $3.75 | |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $4.50 | $5.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | $5.00 | |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $6.00 | $6.75 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $10.00 | $11.25 |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $11.25 | $12.50 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $20.00 | $22.50 |
Cheapest by tier
Popular head-to-heads
Methodology & freshness
Prices are hand-verified against each provider's official pricing page and stored with a source URL and verification date. "Blended cost" assumes 3 input tokens per 1 output token. The oldest verified price in this dataset is from 2026-04-06; the newest is from 2026-07-30. Legacy models remain listed (toggle above) because they still see real API traffic — each links to its recommended successor.
FAQ
Which LLM API is cheapest?
Among the 63 models we track, the cheapest by blended cost (3 input tokens per 1 output token) is Llama 3.1 8B at $0.058 per million tokens.
Why is output more expensive than input?
Generating a token requires a full forward pass through the model, while processing an input token can be batched and cached. Providers price output 3–6x higher than input to reflect this compute cost.
Do these prices include caching discounts?
No — these are the standard, non-cached list prices published by each provider. Prompt caching (where available) can cut input costs significantly for repeated context.
How often are these prices updated?
We verify prices against each provider's official pricing page and re-check on every model addition. The most recently verified price in this table is dated 2026-07-30; the oldest is 2026-04-06.
What does "blended cost" mean?
Blended cost assumes 3 input tokens for every 1 output token — a rough approximation of typical chat/agent usage — computed as (3 × input price + 1 × output price) / 4.
