Llama 3.3 70B API Pricing — $0.59/M input, $0.79/M output
Stable performance: deep knowledge with high reliability.
Groq decommission Aug 16, 2026 — use Muse Spark 1.1 (Meta)
Speed
Not yet measured — see the speed benchmark leaderboard for models we do track.
Cost at scale
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.06 |
| 1,000,000 | $0.64 |
| 10,000,000 | $6.40 |
| 100,000,000 | $64.00 |
How Llama 3.3 70B compares
Llama 3.1 8B — $0.06/MGPT-OSS 20B — $0.13/MGPT-OSS 120B — $0.26/MLlama 4 Maverick — $0.30/MQwen 3.8 30B — $1.20/MGPT-5 Mini — $0.69/MGemini 3.5 Flash Lite — $0.56/MGemini 3.1 Flash Lite — $0.56/M
Head-to-head comparisons
FAQ
Is Llama 3.3 70B cheaper than GPT-5 Mini?
Llama 3.3 70B costs $0.64/M blended tokens, GPT-5 Mini costs $0.69/M — Llama 3.3 70B is cheaper.
How much does 1 million tokens cost with Llama 3.3 70B?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.64. Pure input costs $0.59/M; pure output costs $0.79/M.
Who provides Llama 3.3 70B?
Llama 3.3 70B is served by Groq. Pricing was last verified on 2026-04-06 against https://console.groq.com/docs/models.
