Quarterly data report · Q3 2026
State of LLM API Pricing and Performance
A reproducible snapshot of verified API pricing and controlled gateway performance. The report is descriptive: it does not declare a universal best model.
Data as of 2026-08-08. Stable URL: https://allaiask.com/reports/llm-state-q3-2026
Key findings
- The pricing table covers 65 models. The cheapest blended price is $0.06/M for Llama 3.1 8B; the highest is $67.50/M for GPT-5.4 Pro. [P]
- The median blended price in the published pricing rows is $1.93/M. [P]
- GPT-OSS 120B (Cerebras) is the fastest measured model at 2450 tokens/sec, while Claude Fable 5 is at 41 tokens/sec: a 59.8× spread within this snapshot. [S]
- Throughput and price are separate dimensions: the speed table publishes TTFT, median throughput, p95 throughput, blended price, and cost per throughput unit so readers can make a workload-specific choice. [P][S]
Accessible tables
These tables are the accessible, machine-readable counterpart to any visual comparison. They contain the full rows used for the findings above; use the downloads for analysis.
Pricing snapshot
| Model | Provider | Input / M | Output / M | Blended / M | Verified |
|---|---|---|---|---|---|
| Llama 3.1 8B | Groq | $0.05 | $0.08 | $0.06 | 2026-04-06 |
| Amazon Nova Micro | Amazon | $0.03 | $0.14 | $0.06 | 2026-06-14 |
| Amazon Nova Lite | Amazon | $0.06 | $0.24 | $0.11 | 2026-06-14 |
| GPT-OSS 20B | Groq | $0.07 | $0.30 | $0.13 | 2026-04-06 |
| GPT-5 Nano | OpenAI | $0.05 | $0.40 | $0.14 | 2026-04-06 |
| Ministral 8B | Mistral | $0.15 | $0.15 | $0.15 | 2026-06-14 |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | $0.17 | 2026-06-14 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 | $0.18 | 2026-04-06 | |
| GPT-4o Mini | OpenAI | $0.15 | $0.60 | $0.26 | 2026-04-06 |
| Grok-3 Mini | xAI | $0.15 | $0.60 | $0.26 | 2026-04-06 |
| GPT-OSS 120B | Groq | $0.15 | $0.60 | $0.26 | 2026-04-06 |
| Mistral Small 3.1 | Mistral | $0.15 | $0.60 | $0.26 | 2026-06-14 |
| Llama 4 Maverick | Groq | $0.20 | $0.60 | $0.30 | 2026-07-10 |
| Codestral | Mistral | $0.30 | $0.90 | $0.45 | 2026-06-14 |
| GPT-OSS 120B (Cerebras) | Cerebras | $0.35 | $0.75 | $0.45 | 2026-06-14 |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | $0.45 | 2026-07-30 |
| GPT-5.4 Nano | OpenAI | $0.20 | $1.25 | $0.46 | 2026-04-06 |
| DeepSeek V4 Pro | DeepSeek | $0.43 | $0.87 | $0.54 | 2026-06-14 |
| Gemini 3.5 Flash Lite | $0.25 | $1.50 | $0.56 | 2026-07-22 | |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 | $0.56 | 2026-04-06 | |
| Llama 3.3 70B | Groq | $0.59 | $0.79 | $0.64 | 2026-04-06 |
| GPT-5 Mini | OpenAI | $0.25 | $2.00 | $0.69 | 2026-04-06 |
| Mistral Medium 3 | Mistral | $0.40 | $2.00 | $0.80 | 2026-06-14 |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.85 | 2026-04-06 | |
| GLM-5.1 | Z.ai | $0.60 | $2.20 | $1.00 | 2026-06-19 |
| Qwen 3.7 Plus | Qwen | $0.80 | $2.00 | $1.10 | 2026-07-23 |
| Qwen 3.8 30B | Groq | $0.60 | $3.00 | $1.20 | 2026-07-10 |
| Qwen 3.6 27B | Groq | $0.60 | $3.00 | $1.20 | 2026-06-19 |
| Amazon Nova Pro | Amazon | $0.80 | $3.20 | $1.40 | 2026-06-14 |
| Grok 4.3 | xAI | $1.25 | $2.50 | $1.56 | 2026-05-19 |
| GPT-5.4 Mini | OpenAI | $0.75 | $4.50 | $1.69 | 2026-04-06 |
| Gemini 3.1 Flash | $0.75 | $4.50 | $1.69 | 2026-04-06 | |
| o3-Mini | OpenAI | $1.10 | $4.40 | $1.93 | 2026-04-06 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $2.00 | 2026-04-06 |
| GLM-5.2 | Z.ai | $1.40 | $4.40 | $2.15 | 2026-06-19 |
| GLM 4.7 (Cerebras) | Cerebras | $2.25 | $2.75 | $2.38 | 2026-06-14 |
| Grok-3 | xAI | $2.00 | $4.00 | $2.50 | 2026-04-06 |
| Qwen 3.8 Max | Qwen | $1.60 | $6.40 | $2.80 | 2026-07-10 |
| Qwen 3.7 Max | Qwen | $1.60 | $6.40 | $2.80 | 2026-07-23 |
| Grok-4.20 Reasoning | xAI | $2.00 | $6.00 | $3.00 | 2026-04-06 |
| Grok-4.20 | xAI | $2.00 | $6.00 | $3.00 | 2026-04-06 |
| Mistral Large 3 | Mistral | $2.00 | $6.00 | $3.00 | 2026-06-14 |
| Gemini 3.6 Flash | $1.50 | $9.00 | $3.38 | 2026-07-22 | |
| Gemini 3.5 Flash | $1.50 | $9.00 | $3.38 | 2026-05-21 | |
| GPT-5 | OpenAI | $1.25 | $10.00 | $3.44 | 2026-04-06 |
| GPT-4.1 | OpenAI | $2.00 | $8.00 | $3.50 | 2026-04-06 |
| GPT-4o | OpenAI | $2.50 | $10.00 | $4.38 | 2026-04-06 |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $4.50 | 2026-07-30 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | 2026-04-06 | |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | $5.63 | 2026-04-06 |
| Claude Sonnet 5 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-08-09 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-04-06 |
| Claude Sonnet 4.5 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-04-06 |
| Claude Sonnet 4 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-04-06 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-06-07 |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-04-06 |
| Claude Opus 4.6 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-04-06 |
| Claude Opus 4.5 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-04-06 |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $11.25 | 2026-07-09 |
| GPT-4 Turbo | OpenAI | $10.00 | $30.00 | $15.00 | 2026-04-06 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $20.00 | 2026-06-24 |
| Claude Opus 5 | Anthropic | $15.00 | $75.00 | $30.00 | 2026-08-09 |
| Claude Opus 4.1 | Anthropic | $15.00 | $75.00 | $30.00 | 2026-04-06 |
| Claude Opus 4 | Anthropic | $15.00 | $75.00 | $30.00 | 2026-04-06 |
| GPT-5.4 Pro | OpenAI | $30.00 | $180.00 | $67.50 | 2026-04-06 |
Performance and cost-per-speed snapshot
| Rank | Model | Provider | TTFT ms | Tokens/sec | p95 tokens/sec | Blended / M | $/M ÷ t/s |
|---|---|---|---|---|---|---|---|
| 1 | GPT-OSS 120B (Cerebras) | Cerebras | 90 | 2450 | 2082.5 | $0.45 | $0.0002 |
| 2 | GLM 4.7 (Cerebras) | Cerebras | 110 | 1980 | 1702.8 | $2.38 | $0.0012 |
| 3 | GPT-OSS 20B | Groq | 140 | 1120 | 974.4 | $0.13 | $0.0001 |
| 4 | GPT-OSS 120B | Groq | 160 | 780 | 639.6 | $0.26 | $0.0003 |
| 5 | Qwen 3.8 30B | Groq | 150 | 690 | 634.8 | $1.20 | $0.0017 |
| 6 | Amazon Nova Micro | Amazon | 220 | 168 | 142.8 | $0.06 | $0.0004 |
| 7 | Gemini 3.5 Flash Lite | 240 | 162 | 129.6 | $0.56 | $0.0035 | |
| 8 | Ministral 8B | Mistral | 210 | 158 | 131.14 | $0.15 | $0.0009 |
| 9 | Claude Haiku 4.5 | Anthropic | 260 | 148 | 118.4 | $2.00 | $0.0135 |
| 10 | DeepSeek V4 Flash | DeepSeek | 280 | 132 | 118.8 | $0.17 | $0.0013 |
| 11 | GPT-5.6 Luna | OpenAI | 300 | 126 | 108.36 | $0.45 | $0.0036 |
| 12 | Mistral Small 3.1 | Mistral | 260 | 121 | 110.11 | $0.26 | $0.0022 |
| 13 | Codestral | Mistral | 270 | 118 | 99.12 | $0.45 | $0.0038 |
| 14 | Gemini 3.6 Flash | 310 | 114 | 93.48 | $3.38 | $0.0296 | |
| 15 | Amazon Nova Lite | Amazon | 250 | 108 | 95.04 | $0.11 | $0.0010 |
| 16 | Grok-4.20 | xAI | 290 | 104 | 91.52 | $3.00 | $0.0288 |
| 17 | Grok 4.3 | xAI | 320 | 98 | 82.32 | $1.56 | $0.0159 |
| 18 | Mistral Medium 3 | Mistral | 320 | 92 | 73.6 | $0.80 | $0.0087 |
| 19 | Qwen 3.7 Plus | Qwen | 340 | 84 | 75.6 | $1.10 | $0.0131 |
| 20 | GPT-5.6 Terra | OpenAI | 380 | 78 | 63.96 | $4.50 | $0.0577 |
| 21 | Claude Sonnet 4.6 | Anthropic | 360 | 76 | 69.92 | $6.00 | $0.0789 |
| 22 | DeepSeek V4 Pro | DeepSeek | 480 | 68 | 62.56 | $0.54 | $0.0080 |
| 23 | Amazon Nova Pro | Amazon | 350 | 64 | 55.04 | $1.40 | $0.0219 |
| 24 | Mistral Large 3 | Mistral | 400 | 61 | 50.02 | $3.00 | $0.0492 |
| 25 | Claude Opus 4.8 | Anthropic | 470 | 58 | 53.36 | $10.00 | $0.1724 |
| 26 | Gemini 3.1 Pro | 420 | 55 | 50.6 | $4.50 | $0.0818 | |
| 27 | Grok-4.20 Reasoning | xAI | 540 | 52 | 46.8 | $3.00 | $0.0577 |
| 28 | Qwen 3.7 Max | Qwen | 460 | 49 | 45.08 | $2.80 | $0.0571 |
| 29 | Qwen 3.8 Max | Qwen | 470 | 47 | 37.6 | $2.80 | $0.0596 |
| 30 | GPT-5.6 Sol | OpenAI | 560 | 44 | 35.64 | $11.25 | $0.2557 |
| 31 | Claude Fable 5 | Anthropic | 610 | 41 | 35.67 | $20.00 | $0.4878 |
Methodology
Pricing is read from the versioned provider registry, converted from $/1K to $/1M, and blended as (3 × input + output) ÷ 4. The performance snapshot uses one fixed prompt, 5 runs per model, measured through the All AI Ask gateway from us-east-1. TTFT is reported separately from throughput. Rankings exclude estimated rows.
Limitations
- List prices and one gateway route do not predict your total bill or end-to-end latency.
- Five runs per model are a snapshot, not a confidence interval or a sustained-load test.
- Provider regions, batching, caching, context length, output length, and model availability can change results.
- No historical quarterly baseline is included in this repository, so “key changes” here means notable findings in this version, not a claimed quarter-over-quarter delta.
Download and cite
Download JSON · Download CSV · Read dataset metadata
Citation: All AI Ask. State of LLM API Pricing and Performance — Q3 2026. Data as of 2026-08-08. https://allaiask.com/reports/llm-state-q3-2026. Cite dataset IDs all-ai-ask-pricing-registry-2026-08-08 and all-ai-ask-speed-benchmark-2026-08-08.
[P] Versioned pricing registry · [S] Versioned speed benchmark
