Llama 3.1 8B API Pricing — $0.05/M input, $0.08/M output
Instant response times for lightweight tasks and tool calls.
Groq decommission Aug 16, 2026 — use Muse Spark 1.1 (Meta)
Speed
Not yet measured — see the speed benchmark leaderboard for models we do track.
Cost at scale
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.0058 |
| 1,000,000 | $0.06 |
| 10,000,000 | $0.58 |
| 100,000,000 | $5.75 |
How Llama 3.1 8B compares
GPT-OSS 20B — $0.13/MGPT-OSS 120B — $0.26/MLlama 4 Maverick — $0.30/MLlama 3.3 70B — $0.64/MQwen 3.8 30B — $1.20/MAmazon Nova Micro — $0.06/MAmazon Nova Lite — $0.11/MGPT-OSS 20B — $0.13/M
FAQ
Is Llama 3.1 8B cheaper than Amazon Nova Micro?
Llama 3.1 8B costs $0.06/M blended tokens, Amazon Nova Micro costs $0.06/M — Llama 3.1 8B is cheaper.
How much does 1 million tokens cost with Llama 3.1 8B?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.06. Pure input costs $0.05/M; pure output costs $0.08/M.
Who provides Llama 3.1 8B?
Llama 3.1 8B is served by Groq. Pricing was last verified on 2026-04-06 against https://console.groq.com/docs/models.
