← All tasks

Best LLM for Summarization in 2026

For summarization, Gemini 3.5 Flash Lite is our pick: $0.27/M tokens on a Document summarization workload, 162 tokens/sec, 1M context.

Summarization is a long-input, short-output workload — you pay mostly for the input tokens, so a model's input price and context window matter more here than raw output quality.

What is the best LLM for summarization?

Gemini 3.5 Flash Lite, from Google, is the best fit for summarization at $0.27 per million task tokens on a Document summarization workload, measured at 162 tokens/sec, with a 1M-token context window. No cheaper value pick beats it for this task.

Verified 2026-08-08
Best overall
Gemini 3.5 Flash Lite
Google · $0.27/M
Fit 64/100 — the top requirements match for this task.
Best value
Gemini 3.5 Flash Lite
Google · $0.27/M
The strongest fit among budget and mid-tier priced models.
Fastest
GPT-OSS 120B (Cerebras)
Cerebras · $0.36/M
2450 tokens/sec measured.
Longest context
Gemini 3.1 Pro
Google · $2.16/M
2M token context window.

Evidence

We have not run a controlled test for summarization. This ranking is a requirements match on price, measured throughput, and context window — not a quality comparison. Models that fit this task's requirements are ranked; which one performs best on your prompt is a question you should answer by running it. Run all three side by side →

Ranked — top 8 eligible models

"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.

#ModelProviderFitEvidenceTask price/MTokens/secContextScored on
1Gemini 3.5 Flash LiteGoogle64$0.271621Mprice, context, speed
2Gemini 3.1 ProGoogle61$2.16552Mprice, context, speed
3GLM-5.2Z.ai57$1.451Mprice, context
4Grok 4.3xAI53$1.27981Mprice, context, speed
5Amazon Nova LiteAmazon52$0.06108300Kprice, context, speed
6Gemini 3.6 FlashGoogle51$1.621141Mprice, context, speed
7Grok-4.20xAI49$2.061041Mprice, context, speed
8Grok-4.20 ReasoningxAI49$2.06521Mprice, context, speed

What this costs you

At 10,000 document summarization calls/month:

ModelTask price/MEst. monthly cost
Gemini 3.5 Flash Lite$0.27$137.00
Gemini 3.1 Pro$2.16$1096.00
GLM-5.2$1.45$735.20

How we ranked this

Weights: evidence 0%, price 40%, speed 10%, context 50%.

Requirements: ≥128K context. 32 models eligible.

Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.

Prices verified 2026-08-08.

Related

Google provider hubGemini 3.5 Flash Lite pricingBest LLM for CodingBest LLM for Math & ReasoningBest LLM for Chatbots & Support

FAQ

Why does price dominate this ranking?

Summarization workloads are input-heavy by nature — with a 50K-token document and an 800-token summary, input tokens are over 98% of the bill.

Does a cheaper model summarize worse?

We have not graded this task directly — run your own documents through the top picks below and compare before committing to one at volume.

Should I chunk long documents instead of using a big context window?

Chunking adds complexity and can lose cross-section context; a single large-context call is simpler when the model supports it and the price difference is small.

Run this exact prompt against the top 3

Don't take a ranking's word for it — try Gemini 3.5 Flash Lite and its closest alternatives on your own prompt.

Try It Free