How Much Does RAG Question Answering Cost per Month?
At production volume (60,000 calls/month), the cheapest effective option is Amazon Nova Micro at $27.75/month. The most expensive frontier option, GPT-5.4 Pro, runs $26,093/month — RAG pays mostly for the retrieved context you stuff into the prompt — the answer itself is usually short relative to the chunks that justify it.
How much does rag question answering cost per month?
At production volume (60,000 calls/month), the cheapest effective option for rag question answering is Amazon Nova Micro at $27.75 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $26,093 per month for the same workload.
Token shape
| Shape | large retrieved context in, short answer out |
| Input / output tokens per call | 12K in / 0K out |
| Cacheable input | 50% |
| Batch-eligible | No |
Input is a system prompt plus several retrieved chunks (roughly 12K tokens); a stable system prompt and frequently-reused chunks make about half the input cacheable in a well-tuned pipeline.
Volume
Ranked cost — Production volume
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $28.56 | $16.41 | 0.76× | — |
| Llama 3.1 8Bbudgetlegacy | Groq | $37.92 | $37.48 | 0.77× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $45.60 | $29.40 | — | — |
| Amazon Nova Litebudget | Amazon | $48.96 | $29.06 | 0.92× | — |
| GPT-OSS 20Bbudget | Groq | $61.20 | $75.10 | 2.93× | — |
| Gemini 2.5 Flash Litebudgetlegacy | $81.60 | $49.20 | — | — | |
| Ministral 8Bbudget | Mistral | $112 | $111 | 0.93× | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $108 | $72.91 | 2.60× | ▼1 |
| Mistral Small 3.1budget | Mistral | $122 | $120 | 0.85× | ▲3 |
| GPT-4o Minibudgetlegacy | OpenAI | $122 | $73.80 | — | ▼1 |
| Grok-3 Minibudgetlegacy | xAI | $122 | $122 | — | ▼1 |
| GPT-OSS 120Bbudget | Groq | $122 | $134 | 1.83× | ▼1 |
| Llama 4 Maverickbudgetlegacy | Groq | $158 | $158 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $174 | $103 | 0.79× | ▲1 |
| GPT-5.6 Lunabudget | OpenAI | $173 | $108 | — | ▼1 |
Show all 63 models
| Gemini 3.1 Flash Litebudgetlegacy | $216 | $131 | 0.88× | ▲1 | |
| Gemini 3.5 Flash Litebudget | $216 | $135 | — | ▼1 | |
| GPT-5 Minibudgetlegacy | OpenAI | $228 | $147 | — | — |
| Codestralbudget | Mistral | $238 | $233 | 0.79× | — |
| Gemini 2.5 Flashbudgetlegacy | $276 | $179 | — | ▲1 | |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $270 | $294 | 2.33× | ▼1 |
| Mistral Medium 3budget | Mistral | $336 | $328 | 0.84× | ▲1 |
| DeepSeek V4 Probudget | DeepSeek | $334 | $241 | 3.31× | ▼1 |
| Llama 3.3 70Bbudgetlegacy | Groq | $444 | $440 | 0.81× | — |
| GLM-5.1midlegacy | Z.ai | $485 | $485 | — | — |
| Qwen 3.8 30Bmid | Groq | $504 | $504 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $504 | $504 | — | — |
| Qwen 3.7 Plusmid | Qwen | $624 | $624 | — | — |
| GPT-5.4 Minimidlegacy | OpenAI | $648 | $405 | — | — |
| Gemini 3.1 Flashmidlegacy | $648 | $405 | — | — | |
| Amazon Nova Promid | Amazon | $653 | $394 | — | — |
| Claude Haiku 4.5mid | Anthropic | $840 | $516 | — | — |
| o3-Minimidlegacy | OpenAI | $898 | $541 | — | — |
| Grok 4.3mid | xAI | $960 | $985 | 1.42× | — |
| GPT-5midlegacy | OpenAI | $1,140 | $735 | — | ▲1 |
| Gemini 3.6 Flashmid | $1,296 | $810 | — | ▲1 | |
| Gemini 3.5 Flashmidlegacy | $1,296 | $810 | — | ▲1 | |
| Qwen 3.8 Maxmid | Qwen | $1,306 | $1,306 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $1,306 | $1,306 | — | ▲1 |
| GLM-5.2mid | Z.ai | $1,114 | $1,417 | 3.87× | ▼5 |
| Grok-3midlegacy | xAI | $1,536 | $1,536 | — | — |
| Grok-4.20 Reasoningmid | xAI | $1,584 | $1,584 | — | — |
| Grok-4.20mid | xAI | $1,584 | $1,584 | — | — |
| Mistral Large 3mid | Mistral | $1,584 | $1,584 | — | — |
| Gemini 3.1 Promid | $1,728 | $973 | 0.63× | ▲3 | |
| GPT-4.1midlegacy | OpenAI | $1,632 | $984 | — | ▼1 |
| GPT-5.6 Terramid | OpenAI | $1,728 | $1,080 | — | — |
| GPT-4omidlegacy | OpenAI | $2,040 | $1,230 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,686 | $2,118 | 7.55× | ▼3 |
| GPT-5.4midlegacy | OpenAI | $2,160 | $1,350 | — | — |
| Claude Sonnet 4.6mid | Anthropic | $2,520 | $1,548 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,520 | $1,548 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $2,520 | $1,548 | — | — |
| Claude Opus 4.8mid | Anthropic | $4,200 | $2,556 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $4,200 | $2,580 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $4,200 | $2,580 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $4,200 | $2,580 | — | — |
| GPT-5.6 Solfrontier | OpenAI | $4,320 | $2,700 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $7,920 | $4,680 | — | — |
| Claude Fable 5frontier | Anthropic | $8,400 | $5,160 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $12,600 | $7,740 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $12,600 | $7,740 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $25,920 | $16,373 | 1.04× | — |
Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →
