Best LLM for Translation in 2026
For translation, GPT-OSS 120B (Cerebras) is our pick: $0.55/M tokens on a Translation batch workload, 2450 tokens/sec, 131K context.
Translation workloads are typically high-volume and roughly balanced between input and output tokens, so we weight raw price the heaviest here, with speed as the next-biggest lever for real-time use cases.
What is the best LLM for translation?
GPT-OSS 120B (Cerebras), from Cerebras, is the best fit for translation at $0.55 per million task tokens on a Translation batch workload, measured at 2450 tokens/sec, with a 131K-token context window. No cheaper value pick beats it for this task.
Evidence
Ranked — top 8 eligible models
"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.
| # | Model | Provider | Fit | Evidence | Task price/M | Tokens/sec | Context | Scored on |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-OSS 120B (Cerebras) | Cerebras | 69 | — | $0.55 | 2450 | 131K | price, context, speed |
| 2 | GPT-OSS 20B | Groq | 59 | — | $0.19 | 1120 | 131K | price, context, speed |
| 3 | GLM 4.7 (Cerebras) | Cerebras | 52 | — | $2.50 | 1980 | 200K | price, context, speed |
| 4 | Amazon Nova Micro | Amazon | 52 | — | $0.09 | 168 | 128K | price, context, speed |
| 5 | Amazon Nova Lite | Amazon | 51 | — | $0.15 | 108 | 300K | price, context, speed |
| 6 | GPT-OSS 120B | Groq | 48 | — | $0.38 | 780 | 131K | price, context, speed |
| 7 | GLM-5.2 | Z.ai | 48 | — | $2.90 | — | 1M | price, context |
| 8 | Ministral 8B | Mistral | 47 | — | $0.15 | 158 | 131K | price, context, speed |
What this costs you
At 100,000 translation batch calls/month:
| Model | Task price/M | Est. monthly cost |
|---|---|---|
| GPT-OSS 120B (Cerebras) | $0.55 | $110.00 |
| GPT-OSS 20B | $0.19 | $37.50 |
| GLM 4.7 (Cerebras) | $2.50 | $500.00 |
How we ranked this
Weights: evidence 0%, price 50%, speed 35%, context 15%.
Requirements: none — every current model is eligible. 32 models eligible.
Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.
Prices verified 2026-08-08.
Related
FAQ
Do all models translate equally well?
Quality varies meaningfully by language pair and is not something our current evidence suite measures — test your specific language pairs against the top picks before committing.
Is a reasoning model better for translation?
Rarely — translation is closer to pattern transfer than multi-step reasoning, so the extra latency and cost of a reasoning mode usually isn't worth it.
Does context window matter for translation?
Only for very long documents translated in one call — most translation requests are short enough that any current model's window is more than sufficient.
