Quarterly data report · Q3 2026

State of LLM API Pricing and Performance

A reproducible snapshot of verified API pricing and controlled gateway performance. The report is descriptive: it does not declare a universal best model.

Data as of 2026-08-08. Stable URL: https://allaiask.com/reports/llm-state-q3-2026

Key findings

  • The pricing table covers 65 models. The cheapest blended price is $0.06/M for Llama 3.1 8B; the highest is $67.50/M for GPT-5.4 Pro. [P]
  • The median blended price in the published pricing rows is $1.93/M. [P]
  • GPT-OSS 120B (Cerebras) is the fastest measured model at 2450 tokens/sec, while Claude Fable 5 is at 41 tokens/sec: a 59.8× spread within this snapshot. [S]
  • Throughput and price are separate dimensions: the speed table publishes TTFT, median throughput, p95 throughput, blended price, and cost per throughput unit so readers can make a workload-specific choice. [P][S]

Accessible tables

These tables are the accessible, machine-readable counterpart to any visual comparison. They contain the full rows used for the findings above; use the downloads for analysis.

Pricing snapshot

Pricing rows; all values USD per million tokens.
ModelProviderInput / MOutput / MBlended / MVerified
Llama 3.1 8BGroq$0.05$0.08$0.062026-04-06
Amazon Nova MicroAmazon$0.03$0.14$0.062026-06-14
Amazon Nova LiteAmazon$0.06$0.24$0.112026-06-14
GPT-OSS 20BGroq$0.07$0.30$0.132026-04-06
GPT-5 NanoOpenAI$0.05$0.40$0.142026-04-06
Ministral 8BMistral$0.15$0.15$0.152026-06-14
DeepSeek V4 FlashDeepSeek$0.14$0.28$0.172026-06-14
Gemini 2.5 Flash LiteGoogle$0.10$0.40$0.182026-04-06
GPT-4o MiniOpenAI$0.15$0.60$0.262026-04-06
Grok-3 MinixAI$0.15$0.60$0.262026-04-06
GPT-OSS 120BGroq$0.15$0.60$0.262026-04-06
Mistral Small 3.1Mistral$0.15$0.60$0.262026-06-14
Llama 4 MaverickGroq$0.20$0.60$0.302026-07-10
CodestralMistral$0.30$0.90$0.452026-06-14
GPT-OSS 120B (Cerebras)Cerebras$0.35$0.75$0.452026-06-14
GPT-5.6 LunaOpenAI$0.20$1.20$0.452026-07-30
GPT-5.4 NanoOpenAI$0.20$1.25$0.462026-04-06
DeepSeek V4 ProDeepSeek$0.43$0.87$0.542026-06-14
Gemini 3.5 Flash LiteGoogle$0.25$1.50$0.562026-07-22
Gemini 3.1 Flash LiteGoogle$0.25$1.50$0.562026-04-06
Llama 3.3 70BGroq$0.59$0.79$0.642026-04-06
GPT-5 MiniOpenAI$0.25$2.00$0.692026-04-06
Mistral Medium 3Mistral$0.40$2.00$0.802026-06-14
Gemini 2.5 FlashGoogle$0.30$2.50$0.852026-04-06
GLM-5.1Z.ai$0.60$2.20$1.002026-06-19
Qwen 3.7 PlusQwen$0.80$2.00$1.102026-07-23
Qwen 3.8 30BGroq$0.60$3.00$1.202026-07-10
Qwen 3.6 27BGroq$0.60$3.00$1.202026-06-19
Amazon Nova ProAmazon$0.80$3.20$1.402026-06-14
Grok 4.3xAI$1.25$2.50$1.562026-05-19
GPT-5.4 MiniOpenAI$0.75$4.50$1.692026-04-06
Gemini 3.1 FlashGoogle$0.75$4.50$1.692026-04-06
o3-MiniOpenAI$1.10$4.40$1.932026-04-06
Claude Haiku 4.5Anthropic$1.00$5.00$2.002026-04-06
GLM-5.2Z.ai$1.40$4.40$2.152026-06-19
GLM 4.7 (Cerebras)Cerebras$2.25$2.75$2.382026-06-14
Grok-3xAI$2.00$4.00$2.502026-04-06
Qwen 3.8 MaxQwen$1.60$6.40$2.802026-07-10
Qwen 3.7 MaxQwen$1.60$6.40$2.802026-07-23
Grok-4.20 ReasoningxAI$2.00$6.00$3.002026-04-06
Grok-4.20xAI$2.00$6.00$3.002026-04-06
Mistral Large 3Mistral$2.00$6.00$3.002026-06-14
Gemini 3.6 FlashGoogle$1.50$9.00$3.382026-07-22
Gemini 3.5 FlashGoogle$1.50$9.00$3.382026-05-21
GPT-5OpenAI$1.25$10.00$3.442026-04-06
GPT-4.1OpenAI$2.00$8.00$3.502026-04-06
GPT-4oOpenAI$2.50$10.00$4.382026-04-06
GPT-5.6 TerraOpenAI$2.00$12.00$4.502026-07-30
Gemini 3.1 ProGoogle$2.00$12.00$4.502026-04-06
GPT-5.4OpenAI$2.50$15.00$5.632026-04-06
Claude Sonnet 5Anthropic$3.00$15.00$6.002026-08-09
Claude Sonnet 4.6Anthropic$3.00$15.00$6.002026-04-06
Claude Sonnet 4.5Anthropic$3.00$15.00$6.002026-04-06
Claude Sonnet 4Anthropic$3.00$15.00$6.002026-04-06
Claude Opus 4.8Anthropic$5.00$25.00$10.002026-06-07
Claude Opus 4.7Anthropic$5.00$25.00$10.002026-04-06
Claude Opus 4.6Anthropic$5.00$25.00$10.002026-04-06
Claude Opus 4.5Anthropic$5.00$25.00$10.002026-04-06
GPT-5.6 SolOpenAI$5.00$30.00$11.252026-07-09
GPT-4 TurboOpenAI$10.00$30.00$15.002026-04-06
Claude Fable 5Anthropic$10.00$50.00$20.002026-06-24
Claude Opus 5Anthropic$15.00$75.00$30.002026-08-09
Claude Opus 4.1Anthropic$15.00$75.00$30.002026-04-06
Claude Opus 4Anthropic$15.00$75.00$30.002026-04-06
GPT-5.4 ProOpenAI$30.00$180.00$67.502026-04-06

Performance and cost-per-speed snapshot

Measured rows; throughput is median tokens/sec.
RankModelProviderTTFT msTokens/secp95 tokens/secBlended / M$/M ÷ t/s
1GPT-OSS 120B (Cerebras)Cerebras9024502082.5$0.45$0.0002
2GLM 4.7 (Cerebras)Cerebras11019801702.8$2.38$0.0012
3GPT-OSS 20BGroq1401120974.4$0.13$0.0001
4GPT-OSS 120BGroq160780639.6$0.26$0.0003
5Qwen 3.8 30BGroq150690634.8$1.20$0.0017
6Amazon Nova MicroAmazon220168142.8$0.06$0.0004
7Gemini 3.5 Flash LiteGoogle240162129.6$0.56$0.0035
8Ministral 8BMistral210158131.14$0.15$0.0009
9Claude Haiku 4.5Anthropic260148118.4$2.00$0.0135
10DeepSeek V4 FlashDeepSeek280132118.8$0.17$0.0013
11GPT-5.6 LunaOpenAI300126108.36$0.45$0.0036
12Mistral Small 3.1Mistral260121110.11$0.26$0.0022
13CodestralMistral27011899.12$0.45$0.0038
14Gemini 3.6 FlashGoogle31011493.48$3.38$0.0296
15Amazon Nova LiteAmazon25010895.04$0.11$0.0010
16Grok-4.20xAI29010491.52$3.00$0.0288
17Grok 4.3xAI3209882.32$1.56$0.0159
18Mistral Medium 3Mistral3209273.6$0.80$0.0087
19Qwen 3.7 PlusQwen3408475.6$1.10$0.0131
20GPT-5.6 TerraOpenAI3807863.96$4.50$0.0577
21Claude Sonnet 4.6Anthropic3607669.92$6.00$0.0789
22DeepSeek V4 ProDeepSeek4806862.56$0.54$0.0080
23Amazon Nova ProAmazon3506455.04$1.40$0.0219
24Mistral Large 3Mistral4006150.02$3.00$0.0492
25Claude Opus 4.8Anthropic4705853.36$10.00$0.1724
26Gemini 3.1 ProGoogle4205550.6$4.50$0.0818
27Grok-4.20 ReasoningxAI5405246.8$3.00$0.0577
28Qwen 3.7 MaxQwen4604945.08$2.80$0.0571
29Qwen 3.8 MaxQwen4704737.6$2.80$0.0596
30GPT-5.6 SolOpenAI5604435.64$11.25$0.2557
31Claude Fable 5Anthropic6104135.67$20.00$0.4878

Methodology

Pricing is read from the versioned provider registry, converted from $/1K to $/1M, and blended as (3 × input + output) ÷ 4. The performance snapshot uses one fixed prompt, 5 runs per model, measured through the All AI Ask gateway from us-east-1. TTFT is reported separately from throughput. Rankings exclude estimated rows.

Limitations

  • List prices and one gateway route do not predict your total bill or end-to-end latency.
  • Five runs per model are a snapshot, not a confidence interval or a sustained-load test.
  • Provider regions, batching, caching, context length, output length, and model availability can change results.
  • No historical quarterly baseline is included in this repository, so “key changes” here means notable findings in this version, not a claimed quarter-over-quarter delta.

Download and cite

Download JSON · Download CSV · Read dataset metadata

Citation: All AI Ask. State of LLM API Pricing and Performance — Q3 2026. Data as of 2026-08-08. https://allaiask.com/reports/llm-state-q3-2026. Cite dataset IDs all-ai-ask-pricing-registry-2026-08-08 and all-ai-ask-speed-benchmark-2026-08-08.

[P] Versioned pricing registry · [S] Versioned speed benchmark