LLM Cost Calculator — Real Monthly Cost, Not Just Rate

Raw dataset: data.json. Verbosity measured 2026-06-21. Cite this: All AI Ask LLM Verbosity Index and Effective Cost Dataset, retrieved 2026-06-21.

A $/M rate is not your bill. Set your call shape below and see every priced model ranked by effective monthly cost — adjusted for how many output tokens each model actually spends on a job of this size, not its list price alone.

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.60$24.250.76×
Llama 3.1 8BbudgetlegacyGroq$29.60$28.310.77×
Amazon Nova LitebudgetAmazon$45.60$44.260.92×
GPT-5 NanobudgetlegacyOpenAI$52.00$52.00
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$76.001
Ministral 8BbudgetMistral$82.50$81.770.93×1
GPT-OSS 20BbudgetGroq$57.00$97.532.93×2
Mistral Small 3.1budgetMistral$114$1080.85×4
GPT-4o MinibudgetlegacyOpenAI$114$114
Grok-3 MinibudgetlegacyxAI$114$114
DeepSeek V4 FlashbudgetDeepSeek$86.80$1182.60×3
Llama 4 MaverickbudgetlegacyGroq$138$1381
GPT-OSS 120BbudgetGroq$114$1491.83×2
GPT-5.4 NanobudgetlegacyOpenAI$184$1650.79×1
GPT-5.6 LunabudgetOpenAI$180$1801
Show all 63 models
CodestralbudgetMistral$207$1940.79×
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$2120.88×2
Gemini 3.5 Flash LitebudgetGoogle$225$225
GPT-5 MinibudgetlegacyOpenAI$260$2601
GPT-OSS 120B (Cerebras)budgetCerebras$221$2902.33×3
Mistral Medium 3budgetMistral$332$3100.84×2
Gemini 2.5 FlashbudgetlegacyGoogle$319$319
Llama 3.3 70BbudgetlegacyGroq$339$3280.81×1
DeepSeek V4 ProbudgetDeepSeek$270$4103.31×3
GLM-5.1midlegacyZ.ai$442$442
Qwen 3.8 30BmidGroq$498$498
Qwen 3.6 27BmidlegacyGroq$498$498
Qwen 3.7 PlusmidQwen$524$524
Amazon Nova PromidAmazon$608$608
GPT-5.4 MinimidlegacyOpenAI$675$675
Gemini 3.1 FlashmidlegacyGoogle$675$675
Claude Haiku 4.5midAnthropic$830$8301
o3-MinimidlegacyOpenAI$836$8361
Grok 4.3midxAI$775$8491.42×2
Qwen 3.8 MaxmidQwen$1,216$1,2161
Qwen 3.7 MaxmidQwen$1,216$1,2161
Grok-3midlegacyxAI$1,240$1,2401
GPT-5midlegacyOpenAI$1,300$1,3002
Gemini 3.6 FlashmidGoogle$1,350$1,3502
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,3502
Grok-4.20 ReasoningmidxAI$1,380$1,3802
Grok-4.20midxAI$1,380$1,3802
Mistral Large 3midMistral$1,380$1,3802
Gemini 3.1 PromidGoogle$1,800$1,4890.63×4
GPT-4.1midlegacyOpenAI$1,520$1,5201
GPT-5.6 TerramidOpenAI$1,800$1,8001
GLM-5.2midZ.ai$980$1,8643.87×12
GPT-4omidlegacyOpenAI$1,900$1,9001
GPT-5.4midlegacyOpenAI$2,250$2,2501
Claude Sonnet 4.6midAnthropic$2,490$2,4901
Claude Sonnet 4.5midlegacyAnthropic$2,490$2,4901
Claude Sonnet 4midlegacyAnthropic$2,490$2,4901
GLM 4.7 (Cerebras)midCerebras$1,273$2,5337.55×14
Claude Opus 4.8midAnthropic$4,150$4,0800.96×
Claude Opus 4.7midlegacyAnthropic$4,150$4,150
Claude Opus 4.6midlegacyAnthropic$4,150$4,150
Claude Opus 4.5midlegacyAnthropic$4,150$4,150
GPT-5.6 SolfrontierOpenAI$4,500$4,500
GPT-4 TurbofrontierlegacyOpenAI$6,900$6,900
Claude Fable 5frontierAnthropic$8,300$8,300
Claude Opus 4.1frontierlegacyAnthropic$12,450$12,450
Claude Opus 4frontierlegacyAnthropic$12,450$12,450
GPT-5.4 ProfrontierlegacyOpenAI$27,000$27,5041.04×

List price lied to you

These models move the most once verbosity is priced in — a model that talks more costs more, regardless of its list rate.

  • GLM 4.7 (Cerebras) is priced #39 by list rate but #53 once its 7.55× verbosity is billed — $1,273 list vs $2,533 effective.
  • GLM-5.2 is priced #35 by list rate but #47 once its 3.87× verbosity is billed — $980 list vs $1,864 effective.
  • Mistral Small 3.1 is priced #12 by list rate but #8 once its 0.85× verbosity is billed — $114 list vs $108 effective.
  • Gemini 3.1 Pro is priced #48 by list rate but #44 once its 0.63× verbosity is billed — $1,800 list vs $1,489 effective.
  • DeepSeek V4 Flash is priced #8 by list rate but #11 once its 2.60× verbosity is billed — $86.80 list vs $118 effective.
  • GPT-OSS 120B (Cerebras) is priced #17 by list rate but #20 once its 2.33× verbosity is billed — $221 list vs $290 effective.

The formula, published

outputTokensBilled = outputTokens × (verbosityIndex ?? 1)
inputCost  = inputPerM  × inputTokens        / 1e6
outputCost = outputPerM × outputTokensBilled / 1e6
effective  = (inputCost + outputCost) × callsPerMonth

with caching: inputCost × (1 − cacheableInputPct × 0.9)
with batching: (inputCost + outputCost) × (1 − batchDiscountPct/100)

90% cache-read saving is a stated modelling assumption, not a per-provider sourced figure — providers publish cache-read discounts between 75% and 90% off. The free-form estimator above assumes 30% of input is cache-eligible; pick a workload preset below for a shape-specific assumption instead.

Coverage: 21 of 63 priced models have a measured verbosity index (2+ graded runs). The rest render with a verbosity of — and are shown at unadjusted list price.

Cost by workload shape

LLM Chatbot
growing conversation history in, short reply out
RAG Question Answering
large retrieved context in, short answer out
Coding Agent
large file context in, large diff out
Document Extraction
medium document in, tiny structured JSON out
Long-Document Summarization
very large document in, medium summary out
Content Generation
tiny brief in, long piece out
Classification at Volume
tiny input in, single-label output out
Agentic Tool Loop
many small round-trips per completed task