GPT-OSS 120B
Self-hostable or Groq-speed agentic workflows on open weights.
What are GPT-OSS 120B's specs and price?
GPT-OSS 120B, built by Groq, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.26 per million blended tokens, the 6th-cheapest of 32 models we track.
Specs
| Context window | 131K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2025-08 |
| Knowledge cutoff | 2025-05 |
| Provider | Groq |
Verified 2026-08-08 — source.
Where it ranks
Strengths
- Open-weight MoE flagship
- Ultra-fast on Groq LPUs
- Strong agentic tool-use
More about GPT-OSS 120B
FAQ
What is GPT-OSS 120B's context window?
GPT-OSS 120B has a 131K-token context window and a 33K-token max output — the 22nd-largest context of the 32 current models we track. Source: https://console.groq.com/docs/models, verified 2026-08-08.
Does GPT-OSS 120B support vision or audio input?
No — GPT-OSS 120B is text-only as of 2026-08-08.
Does GPT-OSS 120B have a reasoning or extended-thinking mode?
Yes — GPT-OSS 120B exposes a dedicated reasoning mode for multi-step problems.
When was GPT-OSS 120B released, and what is its knowledge cutoff?
GPT-OSS 120B was released 2025-08 with a knowledge cutoff of 2025-05.
How much does GPT-OSS 120B cost, and who provides it?
GPT-OSS 120B is served by Groq at $0.26/M blended tokens (3:1 input:output) — the 6th-cheapest of 32 current models. Full pricing breakdown: /llm-api-pricing/gpt-oss-120b.
