How Much Does LLM Classification at Volume Cost per Month?

At production volume (5,000,000 calls/month), the cheapest effective option is Amazon Nova Micro at $98.14/month. The most expensive frontier option, GPT-5.4 Pro, runs $93,720/month — Classification is the smallest per-call shape in this cluster, but volume is the entire story — a tiny per-call cost still adds up at millions of calls a month.

How much does classification at volume cost per month?

At production volume (5,000,000 calls/month), the cheapest effective option for classification at volume is Amazon Nova Micro at $98.14 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $93,720 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only.

Token shape

Shapetiny input in, single-label output out
Input / output tokens per call1K in / 0K out
Cacheable input40%
Batch-eligibleYes

Input is one short text plus a stable instruction/label-set prompt; output is a single label or short code. The instruction portion is highly cacheable, and this is the clearest batch-processing candidate in the cluster.

Volume

Side project
500,000 calls/mo
$9.81/mo cheapest
Production
5,000,000 calls/mo
$98.14/mo cheapest
Scale
50,000,000 calls/mo
$981/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$101$66.640.76×
Llama 3.1 8BbudgetlegacyGroq$133$65.580.77×
GPT-5 NanobudgetlegacyOpenAI$165$120
Amazon Nova LitebudgetAmazon$174$1180.92×
GPT-OSS 20BbudgetGroq$218$1382.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$290$200
Ministral 8BbudgetMistral$390$1940.93×1
DeepSeek V4 FlashbudgetDeepSeek$378$2972.60×1
Mistral Small 3.1budgetMistral$435$2130.85×3
GPT-4o MinibudgetlegacyOpenAI$435$3001
Grok-3 MinibudgetlegacyxAI$435$4351
GPT-OSS 120BbudgetGroq$435$2421.83×1
Llama 4 MaverickbudgetlegacyGroq$560$280
GPT-5.4 NanobudgetlegacyOpenAI$625$4190.79×1
GPT-5.6 LunabudgetOpenAI$620$4401
Show all 63 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$775$5320.88×1
Gemini 3.5 Flash LitebudgetGoogle$775$5501
CodestralbudgetMistral$840$4110.79×1
GPT-5 MinibudgetlegacyOpenAI$825$6001
Gemini 2.5 FlashbudgetlegacyGoogle$1,000$7301
GPT-OSS 120B (Cerebras)budgetCerebras$950$1,0502.33×1
Mistral Medium 3budgetMistral$1,200$5840.84×1
DeepSeek V4 ProbudgetDeepSeek$1,175$9843.31×1
Llama 3.3 70BbudgetlegacyGroq$1,554$7690.81×
GLM-5.1midlegacyZ.ai$1,720$1,720
Qwen 3.8 30BmidGroq$1,800$900
Qwen 3.6 27BmidlegacyGroq$1,800$900
Qwen 3.7 PlusmidQwen$2,200$2,200
Amazon Nova PromidAmazon$2,320$1,600
GPT-5.4 MinimidlegacyOpenAI$2,325$1,650
Gemini 3.1 FlashmidlegacyGoogle$2,325$1,650
Claude Haiku 4.5midAnthropic$3,000$2,100
o3-MinimidlegacyOpenAI$3,190$2,200
Grok 4.3midxAI$3,375$3,4801.42×
GPT-5midlegacyOpenAI$4,125$3,0001
Qwen 3.8 MaxmidQwen$4,640$4,6401
Qwen 3.7 MaxmidQwen$4,640$4,6401
Gemini 3.6 FlashmidGoogle$4,650$3,3001
Gemini 3.5 FlashmidlegacyGoogle$4,650$3,3001
GLM-5.2midZ.ai$3,940$5,2033.87×5
Grok-3midlegacyxAI$5,400$5,400
Grok-4.20 ReasoningmidxAI$5,600$5,600
Grok-4.20midxAI$5,600$5,600
Mistral Large 3midMistral$5,600$2,800
Gemini 3.1 PromidGoogle$6,200$3,9560.63×3
GPT-4.1midlegacyOpenAI$5,800$4,0001
GPT-5.6 TerramidOpenAI$6,200$4,400
GPT-4omidlegacyOpenAI$7,250$5,0001
GLM 4.7 (Cerebras)midCerebras$5,900$7,7017.55×3
GPT-5.4midlegacyOpenAI$7,750$5,500
Claude Sonnet 4.6midAnthropic$9,000$6,300
Claude Sonnet 4.5midlegacyAnthropic$9,000$6,300
Claude Sonnet 4midlegacyAnthropic$9,000$6,300
Claude Opus 4.8midAnthropic$15,000$10,4000.96×
Claude Opus 4.7midlegacyAnthropic$15,000$10,500
Claude Opus 4.6midlegacyAnthropic$15,000$10,500
Claude Opus 4.5midlegacyAnthropic$15,000$10,500
GPT-5.6 SolfrontierOpenAI$15,500$11,000
GPT-4 TurbofrontierlegacyOpenAI$28,000$19,000
Claude Fable 5frontierAnthropic$30,000$21,000
Claude Opus 4.1frontierlegacyAnthropic$45,000$31,500
Claude Opus 4frontierlegacyAnthropic$45,000$31,500
GPT-5.4 ProfrontierlegacyOpenAI$93,000$66,7201.04×

Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroLLM Chatbot cost →RAG Question Answering cost →

FAQ

Why does a tiny per-call cost matter here?
At tens of millions of calls a month, a fraction-of-a-cent difference per call compounds into a large monthly delta between models.
Should this always run through a batch API?
Whenever the classification does not need a synchronous response, yes — batch discounts apply directly to a workload this uniform.