How Much Does Document Extraction Cost per Month?

At production volume (500,000 calls/month), the cheapest effective option is Amazon Nova Micro at $118/month. The most expensive frontier option, GPT-5.4 Pro, runs $113,400/month — Extraction reads a medium-sized document and returns a small, fixed-schema JSON object — the output is small and predictable by design.

How much does document extraction cost per month?

At production volume (500,000 calls/month), the cheapest effective option for document extraction is Amazon Nova Micro at $118 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $113,400 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapemedium document in, tiny structured JSON out
Input / output tokens per call6K in / 0K out
Cacheable input15%
Batch-eligibleYes

Input is one document per call (schema and instructions are the only stable, cacheable part); output is a compact JSON object matching a fixed schema. This runs at real volume, so batch discounts matter.

Volume

Side project
50,000 calls/mo
$11.83/mo cheapest
Production
500,000 calls/mo
$118/mo cheapest
Scale
5,000,000 calls/mo
$1,183/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$123$1040.76×
Llama 3.1 8BbudgetlegacyGroq$160$78.850.77×
GPT-5 NanobudgetlegacyOpenAI$200$180
Amazon Nova LitebudgetAmazon$210$1830.92×
GPT-OSS 20BbudgetGroq$263$1672.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$350$310
Ministral 8BbudgetMistral$469$2340.93×1
DeepSeek V4 FlashbudgetDeepSeek$455$4542.60×1
Mistral Small 3.1budgetMistral$525$2570.85×3
GPT-4o MinibudgetlegacyOpenAI$525$4641
Grok-3 MinibudgetlegacyxAI$525$5251
GPT-OSS 120BbudgetGroq$525$2941.83×1
Llama 4 MaverickbudgetlegacyGroq$675$337
GPT-5.4 NanobudgetlegacyOpenAI$756$6420.79×1
GPT-5.6 LunabudgetOpenAI$750$6691
Show all 63 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$938$8140.88×1
Gemini 3.5 Flash LitebudgetGoogle$938$8361
CodestralbudgetMistral$1,013$4940.79×1
GPT-5 MinibudgetlegacyOpenAI$1,000$8991
Gemini 2.5 FlashbudgetlegacyGoogle$1,213$1,0911
GPT-OSS 120B (Cerebras)budgetCerebras$1,144$1,2682.33×1
Mistral Medium 3budgetMistral$1,450$7050.84×1
DeepSeek V4 ProbudgetDeepSeek$1,414$1,4893.31×1
Llama 3.3 70BbudgetlegacyGroq$1,869$9250.81×
GLM-5.1midlegacyZ.ai$2,075$2,075
Qwen 3.8 30BmidGroq$2,175$1,088
Qwen 3.6 27BmidlegacyGroq$2,175$1,088
Qwen 3.7 PlusmidQwen$2,650$2,650
Amazon Nova PromidAmazon$2,800$2,476
GPT-5.4 MinimidlegacyOpenAI$2,813$2,509
Gemini 3.1 FlashmidlegacyGoogle$2,813$2,509
Claude Haiku 4.5midAnthropic$3,625$3,220
o3-MinimidlegacyOpenAI$3,850$3,405
Grok 4.3midxAI$4,063$4,1941.42×
GPT-5midlegacyOpenAI$5,000$4,4941
Qwen 3.8 MaxmidQwen$5,600$5,6001
Qwen 3.7 MaxmidQwen$5,600$5,6001
Gemini 3.6 FlashmidGoogle$5,625$5,0171
Gemini 3.5 FlashmidlegacyGoogle$5,625$5,0171
GLM-5.2midZ.ai$4,750$6,3293.87×5
Grok-3midlegacyxAI$6,500$6,500
Grok-4.20 ReasoningmidxAI$6,750$6,750
Grok-4.20midxAI$6,750$6,750
Mistral Large 3midMistral$6,750$3,375
Gemini 3.1 PromidGoogle$7,500$6,1350.63×3
GPT-4.1midlegacyOpenAI$7,000$6,1901
GPT-5.6 TerramidOpenAI$7,500$6,690
GPT-4omidlegacyOpenAI$8,750$7,7381
GLM 4.7 (Cerebras)midCerebras$7,094$9,3457.55×3
GPT-5.4midlegacyOpenAI$9,375$8,362
Claude Sonnet 4.6midAnthropic$10,875$9,660
Claude Sonnet 4.5midlegacyAnthropic$10,875$9,660
Claude Sonnet 4midlegacyAnthropic$10,875$9,660
Claude Opus 4.8midAnthropic$18,125$15,9750.96×
Claude Opus 4.7midlegacyAnthropic$18,125$16,100
Claude Opus 4.6midlegacyAnthropic$18,125$16,100
Claude Opus 4.5midlegacyAnthropic$18,125$16,100
GPT-5.6 SolfrontierOpenAI$18,750$16,725
GPT-4 TurbofrontierlegacyOpenAI$33,750$29,700
Claude Fable 5frontierAnthropic$36,250$32,200
Claude Opus 4.1frontierlegacyAnthropic$54,375$48,300
Claude Opus 4frontierlegacyAnthropic$54,375$48,300
GPT-5.4 ProfrontierlegacyOpenAI$112,500$101,2501.04×

Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroLLM Chatbot cost →RAG Question Answering cost →

FAQ

Why is only a small share of input cacheable here?
Unlike a chatbot or RAG pipeline, each document is different — only the extraction instructions and schema repeat across calls.
Is this workload a good fit for batch processing?
Often yes — extraction jobs rarely need a synchronous reply, which makes batch APIs (typically 50% off) a straightforward win at this volume.