LLM Speed Benchmarks — Tokens per Second and Latency, Measured August 2026
The fastest model we measure is GPT-OSS 120B (Cerebras) (Cerebras) at 2450 tokens/sec median throughput. That is a 59.8× spread over the slowest model we track, Claude Fable 5 at 41 t/s. Every number below comes from 5 runs per model against one fixed prompt through our own inference gateway from us-east-1 — see the methodology section for what that measurement does and doesn't capture.
Snapshot measured August 8, 2026. Raw dataset: data.json. Cite this: All AI Ask LLM Speed Benchmark Dataset, retrieved 2026-08-08T14:32:00Z.
See the Q3 2026 pricing and performance report for cross-dataset findings, accessible tables, and citation-ready downloads.
Leaderboard
| Rank ↓ | Model | Provider | TTFT | Tokens / sec | p95 t/s | $ / M / (t/s) | Field (7d) |
|---|---|---|---|---|---|---|---|
| #1 | GPT-OSS 120B (Cerebras) n=5 | Cerebras | 90 ms | 2450 | 2083 | $0.0002 | — |
| #2 | GLM 4.7 (Cerebras) n=5 | Cerebras | 110 ms | 1980 | 1703 | $0.0012 | — |
| #3 | GPT-OSS 20B n=5 | Groq | 140 ms | 1120 | 974 | $0.0001 | — |
| #4 | GPT-OSS 120B n=5 | Groq | 160 ms | 780 | 640 | $0.0003 | — |
| #5 | Qwen 3.8 30B n=5 | Groq | 150 ms | 690 | 635 | $0.0017 | — |
| #6 | Amazon Nova Micro n=5 | Amazon | 220 ms | 168 | 143 | $0.0004 | — |
| #7 | Gemini 3.5 Flash Lite n=5 | 240 ms | 162 | 130 | $0.0052 | — | |
| #8 | Ministral 8B n=5 | Mistral | 210 ms | 158 | 131 | $0.0009 | — |
| #9 | Claude Haiku 4.5 n=5 | Anthropic | 260 ms | 148 | 118 | $0.01 | — |
| #10 | DeepSeek V4 Flash n=5 | DeepSeek | 280 ms | 132 | 119 | $0.0050 | — |
| #11 | GPT-5.6 Luna n=5 | OpenAI | 300 ms | 126 | 108 | $0.02 | — |
| #12 | Mistral Small 3.1 n=5 | Mistral | 260 ms | 121 | 110 | $0.0022 | — |
| #13 | Codestral n=5 | Mistral | 270 ms | 118 | 99 | $0.0038 | — |
| #14 | Gemini 3.6 Flash n=5 | 310 ms | 114 | 93 | $0.03 | — | |
| #15 | Amazon Nova Lite n=5 | Amazon | 250 ms | 108 | 95 | $0.0010 | — |
| #16 | Grok-4.20 n=5 | xAI | 290 ms | 104 | 92 | $0.03 | — |
| #17 | Grok 4.3 n=5 | xAI | 320 ms | 98 | 82 | $0.02 | — |
| #18 | Mistral Medium 3 n=5 | Mistral | 320 ms | 92 | 74 | $0.03 | — |
| #19 | Qwen 3.7 Plus n=5 | Qwen | 340 ms | 84 | 76 | $0.01 | — |
| #20 | GPT-5.6 Terra n=5 | OpenAI | 380 ms | 78 | 64 | $0.07 | — |
| #21 | Claude Sonnet 4.6 n=5 | Anthropic | 360 ms | 76 | 70 | $0.08 | — |
| #22 | DeepSeek V4 Pro n=5 | DeepSeek | 480 ms | 68 | 63 | $0.03 | — |
| #23 | Amazon Nova Pro n=5 | Amazon | 350 ms | 64 | 55 | $0.02 | — |
| #24 | Mistral Large 3 n=5 | Mistral | 400 ms | 61 | 50 | $0.01 | — |
| #25 | Claude Opus 4.8 n=5 | Anthropic | 470 ms | 58 | 53 | $0.17 | — |
| #26 | Gemini 3.1 Pro n=5 | 420 ms | 55 | 51 | $0.08 | — | |
| #27 | Grok-4.20 Reasoning n=5 | xAI | 540 ms | 52 | 47 | $0.06 | — |
| #28 | Qwen 3.7 Max n=5 | Qwen | 460 ms | 49 | 45 | $0.06 | — |
| #29 | Qwen 3.8 Max n=5 | Qwen | 470 ms | 47 | 38 | $0.06 | — |
| #30 | GPT-5.6 Sol n=5 | OpenAI | 560 ms | 44 | 36 | $0.26 | — |
| #31 | Claude Fable 5 n=5 | Anthropic | 610 ms | 41 | 36 | $0.49 | — |
Speed vs. cost
Throughput alone doesn't tell you what speed costs. Each point plots a model's tokens/sec against blended $/M ÷ tokens/sec — dollars per million blended tokens divided by throughput, i.e. what you pay for that speed. Lower and further right is better value.
Fastest by tier
Methodology
- One fixed 500-word essay prompt, capped at 1,000 output tokens, sent to every model.
- 5 runs per model; the table reports the median tokens/sec, not the mean, plus the p95 (worst-case-but-one) throughput.
- Time-to-first-token (TTFT) is measured separately from streaming throughput — a slow-starting model isn't penalized in its t/s figure, and vice versa.
- Run from us-east-1 through the All AI Ask inference gateway — every number includes our own gateway's overhead on top of the provider's raw inference speed. This is a real confound, not a footnote: absolute numbers will differ from calling providers directly.
- Where a provider doesn't report token usage, output tokens are estimated from character count. Those rows are flagged
estimatedand excluded from the ranking. - Models are re-benchmarked and republished before this snapshot turns 45 days old.
Batch 39 · server-rendered decision evidence · verified 2026-08-27
Benchmark distributions and snapshot integrity
Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.
Per-model distribution audit
Formula / rubric: Stable rank requires sample count floor + non-estimated tokens + p50/p95 distribution + failure count + snapshot identity.
Provenance: All AI Ask gateway, US region, 2026-08-13 run wave; five-run minimum for rank eligibility. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
model Abatch39-benchmarks-m1-r1 | n=20; TTFT min/p50/p95/max 180/240/710/1,020 ms; throughput 120/98/74/65 t/s | dispersion 18%; failures 0; estimated=false. | Stable rank: n≥5, estimated=false, p95 present. | ELIGIBLE — distribution complete. |
model Bbatch39-benchmarks-m1-r2 | n=3; TTFT 210/260/—/880; throughput 110/92/—/71 | p95 and failure count absent; rank suppressed. | Median without sample floor cannot receive a stable rank. | Unavailable — p95 and failure count are absent |
model Cbatch39-benchmarks-m1-r3 | n=20; output tokens estimated from characters; TTFT observed | speed values present; estimated=true; excluded from leaderboard. | Estimated token rows remain visible but cannot rank. | EXCLUDED — estimated tokens. |
Module citations: All AI Ask benchmark dataset. All AI Ask evidence registry (verified 2026-08-27).
Repeatability and confound ledger
Formula / rubric: Observed delta = matched warm/cold or wave difference; unknown partitions remain narrowly Unavailable.
Provenance: Run wave, UTC window, gateway region, endpoint, network retries, and snapshot retained per sample. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
warm vs coldbatch39-benchmarks-m2-r1 | model A; 10 cold/10 warm; same prompt hash; gateway us-east | TTFT p50 410 ms cold vs 240 ms warm; delta −170 ms. | Report state-specific distributions, not one blended median. | OBSERVED — warm effect closed. |
gateway/providerbatch39-benchmarks-m2-r2 | same model; gateway us-east; provider endpoint region absent | gateway overhead measured at 82 ms on control; provider-only latency Unavailable — not isolated | Do not present gateway result as provider-native latency. | Unavailable — provider-only latency is not isolated |
retry/time windowbatch39-benchmarks-m2-r3 | 20 samples; 2 network retries; UTC 08:00–09:00; snapshot ID present | retry count and window close; retry-attributed throughput Unavailable — not returned | Exclude retry samples from clean distribution if attribution is missing. | Unavailable — retry-attributed throughput is not returned |
Module citations: All AI Ask benchmark methodology. All AI Ask evidence registry (verified 2026-08-27).
Benchmark change ledger
Formula / rubric: Change classification = model/catalog change OR statistically supported movement; raw dataset and methodology hashes must both match.
Provenance: Current 2026-08-13 and prior eligible snapshot; unsupported movement is not a winner change. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
model Abatch39-benchmarks-m3-r1 | current p50 98 t/s; prior 95 t/s; n=20 each; methodology hash equal | Δ +3.2%; bootstrap interval overlaps zero. | Classify as statistically unsupported movement. | NO DECISION CHANGE — same catalog. |
model Bbatch39-benchmarks-m3-r2 | current model ID changed; prior alias; raw dataset hashes differ | speed delta cannot be separated from catalog change. | Do not attribute to runtime improvement. | Unavailable — speed delta cannot be separated from catalog change |
model Cbatch39-benchmarks-m3-r3 | current n=20; prior n=2; methodology hash changed | current p95 exists; prior comparison is underpowered. | No historical rank delta until both snapshots are eligible. | Unavailable — prior comparison is underpowered |
Module citations: All AI Ask raw benchmark data. All AI Ask evidence registry (verified 2026-08-27).
Run your own prompt and watch the timings live
Every prompt you run in the app shows live per-model latency — not a fixed benchmark prompt, your actual workload.
Try It FreeFAQ
Which LLM is fastest?
As of August 2026, GPT-OSS 120B (Cerebras) (Cerebras) is the fastest model we measure at 2450 tokens/sec median throughput.
What is time to first token (TTFT)?
TTFT is the delay between sending a request and receiving the first output token — it captures latency separately from throughput. A model can have fast TTFT but slow streaming, or the reverse; we report both rather than blending them into one number.
Why is my latency different from these numbers?
These figures are measured with a single fixed 500-word prompt through the All AI Ask gateway from one region. Your prompt length, output length, region, and time of day all affect real latency — the "Field data" column shows actual throughput from live user traffic for comparison.
How often do you re-run this benchmark?
This snapshot was measured August 8, 2026 with 5 runs per model. We re-run and republish before the data turns 45 days old.
Is a faster model always the better choice?
No — throughput is one input among several (cost, context window, reasoning quality). The "$ per M ÷ t/s" column on the leaderboard shows what you pay for that speed, and our versus pages weigh speed alongside price and capability for specific model pairs.
