How Much Does RAG Question Answering Cost per Month?

At production volume (60,000 calls/month), the cheapest effective option is Amazon Nova Micro at $27.75/month. The most expensive frontier option, GPT-5.4 Pro, runs $26,093/month — RAG pays mostly for the retrieved context you stuff into the prompt — the answer itself is usually short relative to the chunks that justify it.

How much does rag question answering cost per month?

At production volume (60,000 calls/month), the cheapest effective option for rag question answering is Amazon Nova Micro at $27.75 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $26,093 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapelarge retrieved context in, short answer out
Input / output tokens per call12K in / 0K out
Cacheable input50%
Batch-eligibleNo

Input is a system prompt plus several retrieved chunks (roughly 12K tokens); a stable system prompt and frequently-reused chunks make about half the input cacheable in a well-tuned pipeline.

Volume

Side project
5,000 calls/mo
$2.31/mo cheapest
Production
60,000 calls/mo
$27.75/mo cheapest
Scale
600,000 calls/mo
$278/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$28.56$16.410.76×
Llama 3.1 8BbudgetlegacyGroq$37.92$37.480.77×
GPT-5 NanobudgetlegacyOpenAI$45.60$29.40
Amazon Nova LitebudgetAmazon$48.96$29.060.92×
GPT-OSS 20BbudgetGroq$61.20$75.102.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$81.60$49.20
Ministral 8BbudgetMistral$112$1110.93×1
DeepSeek V4 FlashbudgetDeepSeek$108$72.912.60×1
Mistral Small 3.1budgetMistral$122$1200.85×3
GPT-4o MinibudgetlegacyOpenAI$122$73.801
Grok-3 MinibudgetlegacyxAI$122$1221
GPT-OSS 120BbudgetGroq$122$1341.83×1
Llama 4 MaverickbudgetlegacyGroq$158$158
GPT-5.4 NanobudgetlegacyOpenAI$174$1030.79×1
GPT-5.6 LunabudgetOpenAI$173$1081
Show all 63 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$216$1310.88×1
Gemini 3.5 Flash LitebudgetGoogle$216$1351
GPT-5 MinibudgetlegacyOpenAI$228$147
CodestralbudgetMistral$238$2330.79×
Gemini 2.5 FlashbudgetlegacyGoogle$276$1791
GPT-OSS 120B (Cerebras)budgetCerebras$270$2942.33×1
Mistral Medium 3budgetMistral$336$3280.84×1
DeepSeek V4 ProbudgetDeepSeek$334$2413.31×1
Llama 3.3 70BbudgetlegacyGroq$444$4400.81×
GLM-5.1midlegacyZ.ai$485$485
Qwen 3.8 30BmidGroq$504$504
Qwen 3.6 27BmidlegacyGroq$504$504
Qwen 3.7 PlusmidQwen$624$624
GPT-5.4 MinimidlegacyOpenAI$648$405
Gemini 3.1 FlashmidlegacyGoogle$648$405
Amazon Nova PromidAmazon$653$394
Claude Haiku 4.5midAnthropic$840$516
o3-MinimidlegacyOpenAI$898$541
Grok 4.3midxAI$960$9851.42×
GPT-5midlegacyOpenAI$1,140$7351
Gemini 3.6 FlashmidGoogle$1,296$8101
Gemini 3.5 FlashmidlegacyGoogle$1,296$8101
Qwen 3.8 MaxmidQwen$1,306$1,3061
Qwen 3.7 MaxmidQwen$1,306$1,3061
GLM-5.2midZ.ai$1,114$1,4173.87×5
Grok-3midlegacyxAI$1,536$1,536
Grok-4.20 ReasoningmidxAI$1,584$1,584
Grok-4.20midxAI$1,584$1,584
Mistral Large 3midMistral$1,584$1,584
Gemini 3.1 PromidGoogle$1,728$9730.63×3
GPT-4.1midlegacyOpenAI$1,632$9841
GPT-5.6 TerramidOpenAI$1,728$1,080
GPT-4omidlegacyOpenAI$2,040$1,2301
GLM 4.7 (Cerebras)midCerebras$1,686$2,1187.55×3
GPT-5.4midlegacyOpenAI$2,160$1,350
Claude Sonnet 4.6midAnthropic$2,520$1,548
Claude Sonnet 4.5midlegacyAnthropic$2,520$1,548
Claude Sonnet 4midlegacyAnthropic$2,520$1,548
Claude Opus 4.8midAnthropic$4,200$2,5560.96×
Claude Opus 4.7midlegacyAnthropic$4,200$2,580
Claude Opus 4.6midlegacyAnthropic$4,200$2,580
Claude Opus 4.5midlegacyAnthropic$4,200$2,580
GPT-5.6 SolfrontierOpenAI$4,320$2,700
GPT-4 TurbofrontierlegacyOpenAI$7,920$4,680
Claude Fable 5frontierAnthropic$8,400$5,160
Claude Opus 4.1frontierlegacyAnthropic$12,600$7,740
Claude Opus 4frontierlegacyAnthropic$12,600$7,740
GPT-5.4 ProfrontierlegacyOpenAI$25,920$16,3731.04×

Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroLLM Chatbot cost →Coding Agent cost →

FAQ

Why is RAG input-heavy?
The point of retrieval is to give the model enough grounding text to answer accurately — that context is usually many times larger than the answer it produces.
Does a bigger context window make RAG cheaper?
No — window size is a capacity limit, not a price discount. You still pay per input token regardless of how large the window is.