How Much Does an LLM Chatbot Cost per Month?

At production volume (200,000 calls/month), the cheapest effective option is Amazon Nova Micro at $24.25/month. The most expensive frontier option, GPT-5.4 Pro, runs $27,504/month — A support or product chatbot re-sends a growing conversation history on every turn, so the input side of the bill grows even though each individual reply stays short.

How much does llm chatbot cost per month?

At production volume (200,000 calls/month), the cheapest effective option for llm chatbot is Amazon Nova Micro at $24.25 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $27,504 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapegrowing conversation history in, short reply out
Input / output tokens per call2K in / 0K out
Cacheable input35%
Batch-eligibleNo

Input is a system prompt plus roughly 8-10 turns of history by the middle of a conversation; output is a typical single-paragraph reply. Only the system prompt is stable enough to cache.

Volume

Side project
20,000 calls/mo
$2.42/mo cheapest
Production
200,000 calls/mo
$24.25/mo cheapest
Scale
2,000,000 calls/mo
$242/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.60$18.960.76×
Llama 3.1 8BbudgetlegacyGroq$29.60$28.310.77×
Amazon Nova LitebudgetAmazon$45.60$35.180.92×
GPT-5 NanobudgetlegacyOpenAI$52.00$44.44
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$60.881
Ministral 8BbudgetMistral$82.50$81.770.93×1
GPT-OSS 20BbudgetGroq$57.00$97.532.93×2
Mistral Small 3.1budgetMistral$114$1080.85×4
GPT-4o MinibudgetlegacyOpenAI$114$91.32
Grok-3 MinibudgetlegacyxAI$114$114
DeepSeek V4 FlashbudgetDeepSeek$86.80$96.992.60×3
Llama 4 MaverickbudgetlegacyGroq$138$1381
GPT-OSS 120BbudgetGroq$114$1491.83×2
GPT-5.4 NanobudgetlegacyOpenAI$184$1350.79×1
GPT-5.6 LunabudgetOpenAI$180$1501
Show all 63 models
CodestralbudgetMistral$207$1940.79×
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$1750.88×2
Gemini 3.5 Flash LitebudgetGoogle$225$187
GPT-5 MinibudgetlegacyOpenAI$260$2221
GPT-OSS 120B (Cerebras)budgetCerebras$221$2902.33×3
Mistral Medium 3budgetMistral$332$3100.84×2
Gemini 2.5 FlashbudgetlegacyGoogle$319$274
Llama 3.3 70BbudgetlegacyGroq$339$3280.81×1
DeepSeek V4 ProbudgetDeepSeek$270$3453.31×3
GLM-5.1midlegacyZ.ai$442$442
Qwen 3.8 30BmidGroq$498$498
Qwen 3.6 27BmidlegacyGroq$498$498
Qwen 3.7 PlusmidQwen$524$524
Amazon Nova PromidAmazon$608$487
GPT-5.4 MinimidlegacyOpenAI$675$562
Gemini 3.1 FlashmidlegacyGoogle$675$562
Claude Haiku 4.5midAnthropic$830$6791
o3-MinimidlegacyOpenAI$836$6701
Grok 4.3midxAI$775$8491.42×2
Qwen 3.8 MaxmidQwen$1,216$1,2161
Qwen 3.7 MaxmidQwen$1,216$1,2161
Grok-3midlegacyxAI$1,240$1,2401
GPT-5midlegacyOpenAI$1,300$1,1112
Gemini 3.6 FlashmidGoogle$1,350$1,1232
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,1232
Grok-4.20 ReasoningmidxAI$1,380$1,3802
Grok-4.20midxAI$1,380$1,3802
Mistral Large 3midMistral$1,380$1,3802
Gemini 3.1 PromidGoogle$1,800$1,1870.63×4
GPT-4.1midlegacyOpenAI$1,520$1,2181
GPT-5.6 TerramidOpenAI$1,800$1,4981
GLM-5.2midZ.ai$980$1,8643.87×12
GPT-4omidlegacyOpenAI$1,900$1,5221
GPT-5.4midlegacyOpenAI$2,250$1,8721
Claude Sonnet 4.6midAnthropic$2,490$2,0361
Claude Sonnet 4.5midlegacyAnthropic$2,490$2,0361
Claude Sonnet 4midlegacyAnthropic$2,490$2,0361
GLM 4.7 (Cerebras)midCerebras$1,273$2,5337.55×14
Claude Opus 4.8midAnthropic$4,150$3,3240.96×
Claude Opus 4.7midlegacyAnthropic$4,150$3,394
Claude Opus 4.6midlegacyAnthropic$4,150$3,394
Claude Opus 4.5midlegacyAnthropic$4,150$3,394
GPT-5.6 SolfrontierOpenAI$4,500$3,744
GPT-4 TurbofrontierlegacyOpenAI$6,900$5,388
Claude Fable 5frontierAnthropic$8,300$6,788
Claude Opus 4.1frontierlegacyAnthropic$12,450$10,182
Claude Opus 4frontierlegacyAnthropic$12,450$10,182
GPT-5.4 ProfrontierlegacyOpenAI$27,000$22,9681.04×

Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroRAG Question Answering cost →Coding Agent cost →

FAQ

Why does conversation history dominate the cost?
Most chat APIs are stateless — the full history is re-sent on every turn, so a long conversation costs more per reply even though the model only generates one short answer at a time.
Does prompt caching help chatbot costs?
Only for the stable part of the prompt (system instructions). The conversation history itself changes every turn, so it is rarely cache-eligible.