How Much Does Long-Document Summarization Cost per Month?

At production volume (8,000 calls/month), the cheapest effective option is Amazon Nova Micro at $26.22/month. The most expensive frontier option, GPT-5.4 Pro, runs $23,397/month — Summarization is the most input-heavy shape in this cluster: a large source document in, a medium-length summary out.

How much does long-document summarization cost per month?

At production volume (8,000 calls/month), the cheapest effective option for long-document summarization is Amazon Nova Micro at $26.22 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $23,397 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapevery large document in, medium summary out
Input / output tokens per call90K in / 1K out
Cacheable input10%
Batch-eligibleYes

Input is a full long-form document (report, transcript, contract); output is a structured summary. Each document is usually read once, so little of the input is cache-eligible.

Volume

Side project
800 calls/mo
$2.62/mo cheapest
Production
8,000 calls/mo
$26.22/mo cheapest
Scale
80,000 calls/mo
$262/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.54$23.950.76×
Llama 3.1 8BbudgetlegacyGroq$36.77$18.300.77×
GPT-5 NanobudgetlegacyOpenAI$39.84$36.60
Amazon Nova LitebudgetAmazon$45.50$41.430.92×
GPT-OSS 20BbudgetGroq$56.88$31.222.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$75.84$69.36
DeepSeek V4 FlashbudgetDeepSeek$103$98.722.60×
Ministral 8BbudgetMistral$109$54.670.93×
Mistral Small 3.1budgetMistral$114$56.450.85×3
GPT-4o MinibudgetlegacyOpenAI$114$1041
Grok-3 MinibudgetlegacyxAI$114$1141
GPT-OSS 120BbudgetGroq$114$59.271.83×1
Llama 4 MaverickbudgetlegacyGroq$150$74.88
GPT-5.4 NanobudgetlegacyOpenAI$156$1410.79×1
GPT-5.6 LunabudgetOpenAI$156$1431
Show all 63 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$194$1760.88×1
Gemini 3.5 Flash LitebudgetGoogle$194$1781
GPT-5 MinibudgetlegacyOpenAI$199$183
CodestralbudgetMistral$225$1110.79×
Gemini 2.5 FlashbudgetlegacyGoogle$240$221
GPT-OSS 120B (Cerebras)budgetCerebras$259$2692.33×
Mistral Medium 3budgetMistral$307$1520.84×
DeepSeek V4 ProbudgetDeepSeek$322$3133.31×
Llama 3.3 70BbudgetlegacyGroq$432$2150.81×
GLM-5.1midlegacyZ.ai$453$453
Qwen 3.8 30BmidGroq$461$230
Qwen 3.6 27BmidlegacyGroq$461$230
GPT-5.4 MinimidlegacyOpenAI$583$535
Gemini 3.1 FlashmidlegacyGoogle$583$535
Qwen 3.7 PlusmidQwen$595$595
Amazon Nova PromidAmazon$607$555
Claude Haiku 4.5midAnthropic$768$703
o3-MinimidlegacyOpenAI$834$763
Grok 4.3midxAI$924$9341.42×
GPT-5midlegacyOpenAI$996$915
Gemini 3.6 FlashmidGoogle$1,166$1,0691
Gemini 3.5 FlashmidlegacyGoogle$1,166$1,0691
GLM-5.2midZ.ai$1,050$1,1713.87×2
Qwen 3.8 MaxmidQwen$1,213$1,213
Qwen 3.7 MaxmidQwen$1,213$1,213
Grok-3midlegacyxAI$1,478$1,478
Grok-4.20 ReasoningmidxAI$1,498$1,498
Grok-4.20midxAI$1,498$1,498
Mistral Large 3midMistral$1,498$749
Gemini 3.1 PromidGoogle$1,555$1,3830.63×2
GPT-4.1midlegacyOpenAI$1,517$1,3871
GPT-5.6 TerramidOpenAI$1,555$1,4261
GLM 4.7 (Cerebras)midCerebras$1,646$1,8197.55×
GPT-4omidlegacyOpenAI$1,896$1,734
GPT-5.4midlegacyOpenAI$1,944$1,782
Claude Sonnet 4.6midAnthropic$2,304$2,110
Claude Sonnet 4.5midlegacyAnthropic$2,304$2,110
Claude Sonnet 4midlegacyAnthropic$2,304$2,110
Claude Opus 4.8midAnthropic$3,840$3,5060.96×
Claude Opus 4.7midlegacyAnthropic$3,840$3,516
Claude Opus 4.6midlegacyAnthropic$3,840$3,516
Claude Opus 4.5midlegacyAnthropic$3,840$3,516
GPT-5.6 SolfrontierOpenAI$3,888$3,564
GPT-4 TurbofrontierlegacyOpenAI$7,488$6,840
Claude Fable 5frontierAnthropic$7,680$7,032
Claude Opus 4.1frontierlegacyAnthropic$11,520$10,548
Claude Opus 4frontierlegacyAnthropic$11,520$10,548
GPT-5.4 ProfrontierlegacyOpenAI$23,328$21,4531.04×

Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroLLM Chatbot cost →RAG Question Answering cost →

FAQ

Why does price scale so directly with input tokens here?
At 90,000 input tokens versus 1,200 output tokens, input is roughly 98% of the token volume on every call — the input rate dominates the bill.
Is chunking cheaper than one large call?
Not usually — chunking adds extra model calls and can lose cross-section context, without changing the total tokens processed.