LLM Speed Benchmarks — Tokens per Second and Latency, Measured August 2026

The fastest model we measure is GPT-OSS 120B (Cerebras) (Cerebras) at 2450 tokens/sec median throughput. That is a 59.8× spread over the slowest model we track, Claude Fable 5 at 41 t/s. Every number below comes from 5 runs per model against one fixed prompt through our own inference gateway from us-east-1 — see the methodology section for what that measurement does and doesn't capture.

Snapshot measured August 8, 2026. Raw dataset: data.json. Cite this: All AI Ask LLM Speed Benchmark Dataset, retrieved 2026-08-08T14:32:00Z.

See the Q3 2026 pricing and performance report for cross-dataset findings, accessible tables, and citation-ready downloads.

Leaderboard

RankModelProviderTTFTTokens / secp95 t/s$ / M / (t/s)Field (7d)
#1GPT-OSS 120B (Cerebras)
n=5
Cerebras90 ms24502083$0.0002
#2GLM 4.7 (Cerebras)
n=5
Cerebras110 ms19801703$0.0012
#3GPT-OSS 20B
n=5
Groq140 ms1120974$0.0001
#4GPT-OSS 120B
n=5
Groq160 ms780640$0.0003
#5Qwen 3.8 30B
n=5
Groq150 ms690635$0.0017
#6Amazon Nova Micro
n=5
Amazon220 ms168143$0.0004
#7Gemini 3.5 Flash Lite
n=5
Google240 ms162130$0.0052
#8Ministral 8B
n=5
Mistral210 ms158131$0.0009
#9Claude Haiku 4.5
n=5
Anthropic260 ms148118$0.01
#10DeepSeek V4 Flash
n=5
DeepSeek280 ms132119$0.0050
#11GPT-5.6 Luna
n=5
OpenAI300 ms126108$0.02
#12Mistral Small 3.1
n=5
Mistral260 ms121110$0.0022
#13Codestral
n=5
Mistral270 ms11899$0.0038
#14Gemini 3.6 Flash
n=5
Google310 ms11493$0.03
#15Amazon Nova Lite
n=5
Amazon250 ms10895$0.0010
#16Grok-4.20
n=5
xAI290 ms10492$0.03
#17Grok 4.3
n=5
xAI320 ms9882$0.02
#18Mistral Medium 3
n=5
Mistral320 ms9274$0.03
#19Qwen 3.7 Plus
n=5
Qwen340 ms8476$0.01
#20GPT-5.6 Terra
n=5
OpenAI380 ms7864$0.07
#21Claude Sonnet 4.6
n=5
Anthropic360 ms7670$0.08
#22DeepSeek V4 Pro
n=5
DeepSeek480 ms6863$0.03
#23Amazon Nova Pro
n=5
Amazon350 ms6455$0.02
#24Mistral Large 3
n=5
Mistral400 ms6150$0.01
#25Claude Opus 4.8
n=5
Anthropic470 ms5853$0.17
#26Gemini 3.1 Pro
n=5
Google420 ms5551$0.08
#27Grok-4.20 Reasoning
n=5
xAI540 ms5247$0.06
#28Qwen 3.7 Max
n=5
Qwen460 ms4945$0.06
#29Qwen 3.8 Max
n=5
Qwen470 ms4738$0.06
#30GPT-5.6 Sol
n=5
OpenAI560 ms4436$0.26
#31Claude Fable 5
n=5
Anthropic610 ms4136$0.49

Speed vs. cost

Throughput alone doesn't tell you what speed costs. Each point plots a model's tokens/sec against blended $/M ÷ tokens/sec — dollars per million blended tokens divided by throughput, i.e. what you pay for that speed. Lower and further right is better value.

Tokens per second (log scale, higher = faster) →↑ $ per M blended tokens ÷ t/s (log scale, lower = better value)GPT-OSS 120B (Cerebras) — 2450 t/s, $0.0002/M per t/sGLM 4.7 (Cerebras) — 1980 t/s, $0.0012/M per t/sGPT-OSS 20B — 1120 t/s, $0.0001/M per t/sGPT-OSS 120B — 780 t/s, $0.0003/M per t/sQwen 3.8 30B — 690 t/s, $0.0017/M per t/sAmazon Nova Micro — 168 t/s, $0.0004/M per t/sGemini 3.5 Flash Lite — 162 t/s, $0.0052/M per t/sMinistral 8B — 158 t/s, $0.0009/M per t/sClaude Haiku 4.5 — 148 t/s, $0.0135/M per t/sDeepSeek V4 Flash — 132 t/s, $0.0050/M per t/sGPT-5.6 Luna — 126 t/s, $0.0179/M per t/sMistral Small 3.1 — 121 t/s, $0.0022/M per t/sCodestral — 118 t/s, $0.0038/M per t/sGemini 3.6 Flash — 114 t/s, $0.0263/M per t/sAmazon Nova Lite — 108 t/s, $0.0010/M per t/sGrok-4.20 — 104 t/s, $0.0288/M per t/sGrok 4.3 — 98 t/s, $0.0159/M per t/sMistral Medium 3 — 92 t/s, $0.0326/M per t/sQwen 3.7 Plus — 84 t/s, $0.0131/M per t/sGPT-5.6 Terra — 78 t/s, $0.0721/M per t/sClaude Sonnet 4.6 — 76 t/s, $0.0789/M per t/sDeepSeek V4 Pro — 68 t/s, $0.0291/M per t/sAmazon Nova Pro — 64 t/s, $0.0219/M per t/sMistral Large 3 — 61 t/s, $0.0123/M per t/sClaude Opus 4.8 — 58 t/s, $0.1724/M per t/sGemini 3.1 Pro — 55 t/s, $0.0818/M per t/sGrok-4.20 Reasoning — 52 t/s, $0.0577/M per t/sQwen 3.7 Max — 49 t/s, $0.0571/M per t/sQwen 3.8 Max — 47 t/s, $0.0596/M per t/sGPT-5.6 Sol — 44 t/s, $0.2557/M per t/sClaude Fable 5 — 41 t/s, $0.4878/M per t/s

Fastest by tier

Fastest Budget
GPT-OSS 120B (Cerebras)
2450 t/s · Cerebras
Fastest Mid-tier
GLM 4.7 (Cerebras)
1980 t/s · Cerebras
Fastest Frontier
GPT-5.6 Sol
44 t/s · OpenAI
Full ranking
Fastest AI models
See the dedicated ranking page →

Methodology

  • One fixed 500-word essay prompt, capped at 1,000 output tokens, sent to every model.
  • 5 runs per model; the table reports the median tokens/sec, not the mean, plus the p95 (worst-case-but-one) throughput.
  • Time-to-first-token (TTFT) is measured separately from streaming throughput — a slow-starting model isn't penalized in its t/s figure, and vice versa.
  • Run from us-east-1 through the All AI Ask inference gateway — every number includes our own gateway's overhead on top of the provider's raw inference speed. This is a real confound, not a footnote: absolute numbers will differ from calling providers directly.
  • Where a provider doesn't report token usage, output tokens are estimated from character count. Those rows are flagged estimated and excluded from the ranking.
  • Models are re-benchmarked and republished before this snapshot turns 45 days old.

Batch 39 · server-rendered decision evidence · verified 2026-08-27

Benchmark distributions and snapshot integrity

Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.

Per-model distribution audit

Formula / rubric: Stable rank requires sample count floor + non-estimated tokens + p50/p95 distribution + failure count + snapshot identity.

Provenance: All AI Ask gateway, US region, 2026-08-13 run wave; five-run minimum for rank eligibility. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
model A
batch39-benchmarks-m1-r1
n=20; TTFT min/p50/p95/max 180/240/710/1,020 ms; throughput 120/98/74/65 t/sdispersion 18%; failures 0; estimated=false.Stable rank: n≥5, estimated=false, p95 present.ELIGIBLE — distribution complete.
model B
batch39-benchmarks-m1-r2
n=3; TTFT 210/260/—/880; throughput 110/92/—/71p95 and failure count absent; rank suppressed.Median without sample floor cannot receive a stable rank.Unavailable — p95 and failure count are absent
model C
batch39-benchmarks-m1-r3
n=20; output tokens estimated from characters; TTFT observedspeed values present; estimated=true; excluded from leaderboard.Estimated token rows remain visible but cannot rank.EXCLUDED — estimated tokens.

Module citations: All AI Ask benchmark dataset. All AI Ask evidence registry (verified 2026-08-27).

Repeatability and confound ledger

Formula / rubric: Observed delta = matched warm/cold or wave difference; unknown partitions remain narrowly Unavailable.

Provenance: Run wave, UTC window, gateway region, endpoint, network retries, and snapshot retained per sample. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
warm vs cold
batch39-benchmarks-m2-r1
model A; 10 cold/10 warm; same prompt hash; gateway us-eastTTFT p50 410 ms cold vs 240 ms warm; delta −170 ms.Report state-specific distributions, not one blended median.OBSERVED — warm effect closed.
gateway/provider
batch39-benchmarks-m2-r2
same model; gateway us-east; provider endpoint region absentgateway overhead measured at 82 ms on control; provider-only latency Unavailable — not isolatedDo not present gateway result as provider-native latency.Unavailable — provider-only latency is not isolated
retry/time window
batch39-benchmarks-m2-r3
20 samples; 2 network retries; UTC 08:00–09:00; snapshot ID presentretry count and window close; retry-attributed throughput Unavailable — not returnedExclude retry samples from clean distribution if attribution is missing.Unavailable — retry-attributed throughput is not returned

Module citations: All AI Ask benchmark methodology. All AI Ask evidence registry (verified 2026-08-27).

Benchmark change ledger

Formula / rubric: Change classification = model/catalog change OR statistically supported movement; raw dataset and methodology hashes must both match.

Provenance: Current 2026-08-13 and prior eligible snapshot; unsupported movement is not a winner change. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
model A
batch39-benchmarks-m3-r1
current p50 98 t/s; prior 95 t/s; n=20 each; methodology hash equalΔ +3.2%; bootstrap interval overlaps zero.Classify as statistically unsupported movement.NO DECISION CHANGE — same catalog.
model B
batch39-benchmarks-m3-r2
current model ID changed; prior alias; raw dataset hashes differspeed delta cannot be separated from catalog change.Do not attribute to runtime improvement.Unavailable — speed delta cannot be separated from catalog change
model C
batch39-benchmarks-m3-r3
current n=20; prior n=2; methodology hash changedcurrent p95 exists; prior comparison is underpowered.No historical rank delta until both snapshots are eligible.Unavailable — prior comparison is underpowered

Module citations: All AI Ask raw benchmark data. All AI Ask evidence registry (verified 2026-08-27).

Reproduce the benchmark run

Run your own prompt and watch the timings live

Every prompt you run in the app shows live per-model latency — not a fixed benchmark prompt, your actual workload.

Try It Free

FAQ

Which LLM is fastest?

As of August 2026, GPT-OSS 120B (Cerebras) (Cerebras) is the fastest model we measure at 2450 tokens/sec median throughput.

What is time to first token (TTFT)?

TTFT is the delay between sending a request and receiving the first output token — it captures latency separately from throughput. A model can have fast TTFT but slow streaming, or the reverse; we report both rather than blending them into one number.

Why is my latency different from these numbers?

These figures are measured with a single fixed 500-word prompt through the All AI Ask gateway from one region. Your prompt length, output length, region, and time of day all affect real latency — the "Field data" column shows actual throughput from live user traffic for comparison.

How often do you re-run this benchmark?

This snapshot was measured August 8, 2026 with 5 runs per model. We re-run and republish before the data turns 45 days old.

Is a faster model always the better choice?

No — throughput is one input among several (cost, context window, reasoning quality). The "$ per M ÷ t/s" column on the leaderboard shows what you pay for that speed, and our versus pages weigh speed alongside price and capability for specific model pairs.