Best LLM for Chatbots & Support in 2026
For chatbots & support, GPT-OSS 120B (Cerebras) is our pick: $0.48/M tokens on a Support chat turn workload, 2450 tokens/sec, 131K context.
Support chat is a latency-sensitive, extremely high-volume workload — users notice a slow reply far more than a slightly less polished one, and the per-call cost compounds fast at real support volumes.
What is the best LLM for chatbots & support?
GPT-OSS 120B (Cerebras), from Cerebras, is the best fit for chatbots & support at $0.48 per million task tokens on a Support chat turn workload, measured at 2450 tokens/sec, with a 131K-token context window. No cheaper value pick beats it for this task.
Evidence
Ranked — top 8 eligible models
"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.
| # | Model | Provider | Fit | Evidence | Task price/M | Tokens/sec | Context | Scored on |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-OSS 120B (Cerebras) | Cerebras | 75 | — | $0.48 | 2450 | 131K | price, context, speed |
| 2 | GPT-OSS 20B | Groq | 59 | — | $0.15 | 1120 | 131K | price, context, speed |
| 3 | GLM 4.7 (Cerebras) | Cerebras | 55 | — | $2.42 | 1980 | 200K | price, context, speed |
| 4 | GPT-OSS 120B | Groq | 48 | — | $0.30 | 780 | 131K | price, context, speed |
| 5 | Amazon Nova Micro | Amazon | 47 | — | $0.07 | 168 | 128K | price, context, speed |
| 6 | GLM-5.2 | Z.ai | 46 | — | $2.40 | — | 1M | price, context |
| 7 | Amazon Nova Lite | Amazon | 45 | — | $0.12 | 108 | 300K | price, context, speed |
| 8 | Ministral 8B | Mistral | 41 | — | $0.15 | 158 | 131K | price, context, speed |
What this costs you
At 500,000 support chat turn calls/month:
| Model | Task price/M | Est. monthly cost |
|---|---|---|
| GPT-OSS 120B (Cerebras) | $0.48 | $145.00 |
| GPT-OSS 20B | $0.15 | $45.00 |
| GLM 4.7 (Cerebras) | $2.42 | $725.00 |
How we ranked this
Weights: evidence 0%, price 45%, speed 45%, context 10%.
Requirements: none — every current model is eligible. 32 models eligible.
Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.
Prices verified 2026-08-08.
Related
FAQ
Why not use a flagship model for every support ticket?
At 500,000 turns a month the price gap between a flagship and a fast budget model is enormous, and most support turns don't need frontier-level reasoning.
How fast does a chatbot model need to be?
Aim for sub-second time-to-first-token where possible — measured tokens/sec on this page is a proxy for how quickly a reply starts streaming.
Should I escalate hard questions to a bigger model?
Yes — a common pattern is routing routine turns to a fast, cheap model and escalating low-confidence or complex turns to a stronger one.
