Best LLM for Every Task — 10 Tasks, 32 Models Ranked, Updated August 2026
Raw dataset: data.json. Cite this: All AI Ask LLM Task Recommendation Dataset, retrieved 2026-08-08.
For coding, GLM-5.2 is our pick at $2.00/M tokens. For math & reasoning, GLM-5.2 is our pick at $3.40/M tokens. For chatbots & support, GPT-OSS 120B (Cerebras) is our pick at $0.48/M tokens. Every ranking below shows its formula, requirements, and evidence status — never a bare number.
| Task | Our pick | Task price/M | Evidence | Eligible models | Best budget pick |
|---|---|---|---|---|---|
| Coding | GLM-5.2 | $2.00 | ✓ 16 graded runs | 32 | GLM-5.2 |
| Structured Data Extraction | Amazon Nova Micro | $0.06 | ✓ 13 graded runs | 32 | Amazon Nova Micro |
| Writing & Content | GPT-OSS 120B (Cerebras) | $0.57 | ✓ 13 graded runs | 32 | GPT-OSS 120B (Cerebras) |
| Math & Reasoning | GLM-5.2 | $3.40 | ✓ 3 graded runs | 18 | GLM-5.2 |
| Agents & Tool Use | GLM-5.2 | $2.00 | ✓ 3 graded runs | 18 | GLM-5.2 |
| Long Documents & RAG | Gemini 3.1 Pro | $2.07 | — requirements fit | 21 | Gemini 3.1 Pro |
| Summarization | Gemini 3.5 Flash Lite | $0.27 | — requirements fit | 32 | Gemini 3.5 Flash Lite |
| Chatbots & Support | GPT-OSS 120B (Cerebras) | $0.48 | — requirements fit | 32 | GPT-OSS 120B (Cerebras) |
| Translation | GPT-OSS 120B (Cerebras) | $0.55 | — requirements fit | 32 | GPT-OSS 120B (Cerebras) |
| Image Understanding | Gemini 3.5 Flash Lite | $0.51 | — requirements fit | 22 | Gemini 3.5 Flash Lite |
Data verified 2026-08-08. Every number is derived from our live pricing, speed, and graded-test datasets — see each task page for the exact formula.
How these rankings work
Each task page defines hard eligibility requirements (a minimum context window, vision support, a reasoning mode) and a set of weights across price, measured speed, context window, and — where we have run a graded test — accuracy. Sub-scores are normalised within that task's eligible model set, never globally, and a model missing a measurement is never scored as zero — its weight is redistributed across what we do have. See "How we ranked this" on any task page for the exact numbers.
More data clusters
Don't take a ranking's word for it
Run your own prompt against every pick on this page in one workspace, side by side.
Try It Free