How Much Does Long-Document Summarization Cost per Month?
At production volume (8,000 calls/month), the cheapest effective option is Amazon Nova Micro at $26.22/month. The most expensive frontier option, GPT-5.4 Pro, runs $23,397/month — Summarization is the most input-heavy shape in this cluster: a large source document in, a medium-length summary out.
How much does long-document summarization cost per month?
At production volume (8,000 calls/month), the cheapest effective option for long-document summarization is Amazon Nova Micro at $26.22 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $23,397 per month for the same workload.
Token shape
| Shape | very large document in, medium summary out |
| Input / output tokens per call | 90K in / 1K out |
| Cacheable input | 10% |
| Batch-eligible | Yes |
Input is a full long-form document (report, transcript, contract); output is a structured summary. Each document is usually read once, so little of the input is cache-eligible.
Volume
Ranked cost — Production volume (caching + batch applied where available)
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $26.54 | $23.95 | 0.76× | — |
| Llama 3.1 8Bbudgetlegacy | Groq | $36.77 | $18.30 | 0.77× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $39.84 | $36.60 | — | — |
| Amazon Nova Litebudget | Amazon | $45.50 | $41.43 | 0.92× | — |
| GPT-OSS 20Bbudget | Groq | $56.88 | $31.22 | 2.93× | — |
| Gemini 2.5 Flash Litebudgetlegacy | $75.84 | $69.36 | — | — | |
| DeepSeek V4 Flashbudget | DeepSeek | $103 | $98.72 | 2.60× | — |
| Ministral 8Bbudget | Mistral | $109 | $54.67 | 0.93× | — |
| Mistral Small 3.1budget | Mistral | $114 | $56.45 | 0.85× | ▲3 |
| GPT-4o Minibudgetlegacy | OpenAI | $114 | $104 | — | ▼1 |
| Grok-3 Minibudgetlegacy | xAI | $114 | $114 | — | ▼1 |
| GPT-OSS 120Bbudget | Groq | $114 | $59.27 | 1.83× | ▼1 |
| Llama 4 Maverickbudgetlegacy | Groq | $150 | $74.88 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $156 | $141 | 0.79× | ▲1 |
| GPT-5.6 Lunabudget | OpenAI | $156 | $143 | — | ▼1 |
Show all 63 models
| Gemini 3.1 Flash Litebudgetlegacy | $194 | $176 | 0.88× | ▲1 | |
| Gemini 3.5 Flash Litebudget | $194 | $178 | — | ▼1 | |
| GPT-5 Minibudgetlegacy | OpenAI | $199 | $183 | — | — |
| Codestralbudget | Mistral | $225 | $111 | 0.79× | — |
| Gemini 2.5 Flashbudgetlegacy | $240 | $221 | — | — | |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $259 | $269 | 2.33× | — |
| Mistral Medium 3budget | Mistral | $307 | $152 | 0.84× | — |
| DeepSeek V4 Probudget | DeepSeek | $322 | $313 | 3.31× | — |
| Llama 3.3 70Bbudgetlegacy | Groq | $432 | $215 | 0.81× | — |
| GLM-5.1midlegacy | Z.ai | $453 | $453 | — | — |
| Qwen 3.8 30Bmid | Groq | $461 | $230 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $461 | $230 | — | — |
| GPT-5.4 Minimidlegacy | OpenAI | $583 | $535 | — | — |
| Gemini 3.1 Flashmidlegacy | $583 | $535 | — | — | |
| Qwen 3.7 Plusmid | Qwen | $595 | $595 | — | — |
| Amazon Nova Promid | Amazon | $607 | $555 | — | — |
| Claude Haiku 4.5mid | Anthropic | $768 | $703 | — | — |
| o3-Minimidlegacy | OpenAI | $834 | $763 | — | — |
| Grok 4.3mid | xAI | $924 | $934 | 1.42× | — |
| GPT-5midlegacy | OpenAI | $996 | $915 | — | — |
| Gemini 3.6 Flashmid | $1,166 | $1,069 | — | ▲1 | |
| Gemini 3.5 Flashmidlegacy | $1,166 | $1,069 | — | ▲1 | |
| GLM-5.2mid | Z.ai | $1,050 | $1,171 | 3.87× | ▼2 |
| Qwen 3.8 Maxmid | Qwen | $1,213 | $1,213 | — | — |
| Qwen 3.7 Maxmid | Qwen | $1,213 | $1,213 | — | — |
| Grok-3midlegacy | xAI | $1,478 | $1,478 | — | — |
| Grok-4.20 Reasoningmid | xAI | $1,498 | $1,498 | — | — |
| Grok-4.20mid | xAI | $1,498 | $1,498 | — | — |
| Mistral Large 3mid | Mistral | $1,498 | $749 | — | — |
| Gemini 3.1 Promid | $1,555 | $1,383 | 0.63× | ▲2 | |
| GPT-4.1midlegacy | OpenAI | $1,517 | $1,387 | — | ▼1 |
| GPT-5.6 Terramid | OpenAI | $1,555 | $1,426 | — | ▼1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,646 | $1,819 | 7.55× | — |
| GPT-4omidlegacy | OpenAI | $1,896 | $1,734 | — | — |
| GPT-5.4midlegacy | OpenAI | $1,944 | $1,782 | — | — |
| Claude Sonnet 4.6mid | Anthropic | $2,304 | $2,110 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,304 | $2,110 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $2,304 | $2,110 | — | — |
| Claude Opus 4.8mid | Anthropic | $3,840 | $3,506 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $3,840 | $3,516 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $3,840 | $3,516 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $3,840 | $3,516 | — | — |
| GPT-5.6 Solfrontier | OpenAI | $3,888 | $3,564 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $7,488 | $6,840 | — | — |
| Claude Fable 5frontier | Anthropic | $7,680 | $7,032 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $11,520 | $10,548 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $11,520 | $10,548 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $23,328 | $21,453 | 1.04× | — |
Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →
