Best LLM for Every Task — 10 Tasks, 32 Models Ranked, Updated August 2026

Raw dataset: data.json. Cite this: All AI Ask LLM Task Recommendation Dataset, retrieved 2026-08-08.

For coding, GLM-5.2 is our pick at $2.00/M tokens. For math & reasoning, GLM-5.2 is our pick at $3.40/M tokens. For chatbots & support, GPT-OSS 120B (Cerebras) is our pick at $0.48/M tokens. Every ranking below shows its formula, requirements, and evidence status — never a bare number.

"Fit" is a requirements match, not a quality benchmark. It combines price, measured speed, context window, and — where we have run it — graded accuracy on that exact task. Five of the ten tasks below carry first-party graded evidence; the other five are honest requirements-fit rankings and say so on the page.
TaskOur pickTask price/MEvidenceEligible modelsBest budget pick
CodingGLM-5.2$2.00✓ 16 graded runs32GLM-5.2
Structured Data ExtractionAmazon Nova Micro$0.06✓ 13 graded runs32Amazon Nova Micro
Writing & ContentGPT-OSS 120B (Cerebras)$0.57✓ 13 graded runs32GPT-OSS 120B (Cerebras)
Math & ReasoningGLM-5.2$3.40✓ 3 graded runs18GLM-5.2
Agents & Tool UseGLM-5.2$2.00✓ 3 graded runs18GLM-5.2
Long Documents & RAGGemini 3.1 Pro$2.07— requirements fit21Gemini 3.1 Pro
SummarizationGemini 3.5 Flash Lite$0.27— requirements fit32Gemini 3.5 Flash Lite
Chatbots & SupportGPT-OSS 120B (Cerebras)$0.48— requirements fit32GPT-OSS 120B (Cerebras)
TranslationGPT-OSS 120B (Cerebras)$0.55— requirements fit32GPT-OSS 120B (Cerebras)
Image UnderstandingGemini 3.5 Flash Lite$0.51— requirements fit22Gemini 3.5 Flash Lite

Data verified 2026-08-08. Every number is derived from our live pricing, speed, and graded-test datasets — see each task page for the exact formula.

How these rankings work

Each task page defines hard eligibility requirements (a minimum context window, vision support, a reasoning mode) and a set of weights across price, measured speed, context window, and — where we have run a graded test — accuracy. Sub-scores are normalised within that task's eligible model set, never globally, and a model missing a measurement is never scored as zero — its weight is redistributed across what we do have. See "How we ranked this" on any task page for the exact numbers.

More data clusters

API PricingSpeed BenchmarksProvidersHead-to-HeadModel Deprecations

Don't take a ranking's word for it

Run your own prompt against every pick on this page in one workspace, side by side.

Try It Free