How Much Does a Coding Agent Cost per Month?

At production volume (20,000 calls/month), the cheapest effective option is Amazon Nova Micro at $18.26/month. The most expensive frontier option, GPT-5.4 Pro, runs $19,488/month — A coding agent reads large amounts of file context and writes substantial diffs — both sides of the bill are large compared to a chat turn.

How much does coding agent cost per month?

At production volume (20,000 calls/month), the cheapest effective option for coding agent is Amazon Nova Micro at $18.26 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $19,488 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapelarge file context in, large diff out
Input / output tokens per call20K in / 2K out
Cacheable input60%
Batch-eligibleNo

Input is the files and instructions relevant to one change; output is a multi-file diff plus explanation. Repeated reads of the same files across a session make a majority of input cacheable.

Volume

Side project
2,000 calls/mo
$1.83/mo cheapest
Production
20,000 calls/mo
$18.26/mo cheapest
Scale
200,000 calls/mo
$183/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$19.60$10.700.76×
Llama 3.1 8BbudgetlegacyGroq$23.20$22.460.77×
Amazon Nova LitebudgetAmazon$33.60$19.870.92×
GPT-5 NanobudgetlegacyOpenAI$36.00$25.20
Gemini 2.5 Flash LitebudgetlegacyGoogle$56.00$34.401
GPT-OSS 20BbudgetGroq$42.00$65.162.93×1
Ministral 8BbudgetMistral$66.00$65.580.93×
Mistral Small 3.1budgetMistral$84.00$80.400.85×4
GPT-4o MinibudgetlegacyOpenAI$84.00$51.60
Grok-3 MinibudgetlegacyxAI$84.00$84.00
DeepSeek V4 FlashbudgetDeepSeek$67.20$54.882.60×3
GPT-OSS 120BbudgetGroq$84.00$1041.83×1
Llama 4 MaverickbudgetlegacyGroq$104$104
GPT-5.4 NanobudgetlegacyOpenAI$130$76.300.79×1
GPT-5.6 LunabudgetOpenAI$128$84.801
Show all 63 models
CodestralbudgetMistral$156$1480.79×
Gemini 3.1 Flash LitebudgetlegacyGoogle$160$98.800.88×1
Gemini 3.5 Flash LitebudgetGoogle$160$1061
GPT-5 MinibudgetlegacyOpenAI$180$1261
GPT-OSS 120B (Cerebras)budgetCerebras$170$2102.33×1
Gemini 2.5 FlashbudgetlegacyGoogle$220$1551
Mistral Medium 3budgetMistral$240$2270.84×1
Llama 3.3 70BbudgetlegacyGroq$268$2620.81×1
DeepSeek V4 ProbudgetDeepSeek$209$1953.31×3
GLM-5.1midlegacyZ.ai$328$328
Qwen 3.8 30BmidGroq$360$360
Qwen 3.6 27BmidlegacyGroq$360$360
Qwen 3.7 PlusmidQwen$400$400
Amazon Nova PromidAmazon$448$275
GPT-5.4 MinimidlegacyOpenAI$480$318
Gemini 3.1 FlashmidlegacyGoogle$480$318
Claude Haiku 4.5midAnthropic$600$3841
o3-MinimidlegacyOpenAI$616$3781
Grok 4.3midxAI$600$6421.42×2
Qwen 3.8 MaxmidQwen$896$8961
Qwen 3.7 MaxmidQwen$896$8961
GPT-5midlegacyOpenAI$900$6301
Grok-3midlegacyxAI$960$9601
Gemini 3.6 FlashmidGoogle$960$6361
Gemini 3.5 FlashmidlegacyGoogle$960$6361
Grok-4.20 ReasoningmidxAI$1,040$1,0402
Grok-4.20midxAI$1,040$1,0402
Mistral Large 3midMistral$1,040$1,0402
Gemini 3.1 PromidGoogle$1,280$6700.63×4
GPT-4.1midlegacyOpenAI$1,120$6881
GLM-5.2midZ.ai$736$1,2413.87×11
GPT-5.6 TerramidOpenAI$1,280$848
GPT-4omidlegacyOpenAI$1,400$8601
GPT-5.4midlegacyOpenAI$1,600$1,0601
GLM 4.7 (Cerebras)midCerebras$1,010$1,7307.55×8
Claude Sonnet 4.6midAnthropic$1,800$1,152
Claude Sonnet 4.5midlegacyAnthropic$1,800$1,152
Claude Sonnet 4midlegacyAnthropic$1,800$1,152
Claude Opus 4.8midAnthropic$3,000$1,8800.96×
Claude Opus 4.7midlegacyAnthropic$3,000$1,920
Claude Opus 4.6midlegacyAnthropic$3,000$1,920
Claude Opus 4.5midlegacyAnthropic$3,000$1,920
GPT-5.6 SolfrontierOpenAI$3,200$2,120
GPT-4 TurbofrontierlegacyOpenAI$5,200$3,040
Claude Fable 5frontierAnthropic$6,000$3,840
Claude Opus 4.1frontierlegacyAnthropic$9,000$5,760
Claude Opus 4frontierlegacyAnthropic$9,000$5,760
GPT-5.4 ProfrontierlegacyOpenAI$19,200$13,0081.04×

Prompt caching assumes a 90% saving on the cacheable share of input — a stated modelling assumption, providers publish 75-90% off. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroLLM Chatbot cost →RAG Question Answering cost →

FAQ

Why is output so much larger here than a chatbot?
A code change often touches multiple functions or files, and a coding model typically explains what it changed — both add up to far more output tokens than a chat reply.
Does verbosity matter more for coding agents?
Yes — a model that pads its diffs with unrequested commentary bills for tokens that add no value, and that shows up directly in the effective-cost ranking below.