GPT-OSS 120B (Cerebras)
Agentic workflows where raw inference speed is the deciding factor.
What are GPT-OSS 120B (Cerebras)'s specs and price?
GPT-OSS 120B (Cerebras), built by Cerebras, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.45 per million blended tokens, the 9th-cheapest of 32 models we track.
Specs
| Context window | 131K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2025-08 |
| Knowledge cutoff | 2025-05 |
| Provider | Cerebras |
Verified 2026-08-08 — source.
Where it ranks
Strengths
- Same open-weight flagship as gpt-oss-120b
- Wafer-scale inference at ~5,000+ chars/s
- Fastest hosting option for this model
More about GPT-OSS 120B (Cerebras)
FAQ
What is GPT-OSS 120B (Cerebras)'s context window?
GPT-OSS 120B (Cerebras) has a 131K-token context window and a 33K-token max output — the 29th-largest context of the 32 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-08.
Does GPT-OSS 120B (Cerebras) support vision or audio input?
No — GPT-OSS 120B (Cerebras) is text-only as of 2026-08-08.
Does GPT-OSS 120B (Cerebras) have a reasoning or extended-thinking mode?
Yes — GPT-OSS 120B (Cerebras) exposes a dedicated reasoning mode for multi-step problems.
When was GPT-OSS 120B (Cerebras) released, and what is its knowledge cutoff?
GPT-OSS 120B (Cerebras) was released 2025-08 with a knowledge cutoff of 2025-05.
How much does GPT-OSS 120B (Cerebras) cost, and who provides it?
GPT-OSS 120B (Cerebras) is served by Cerebras at $0.45/M blended tokens (3:1 input:output) — the 9th-cheapest of 32 current models. Full pricing breakdown: /llm-api-pricing/cerebras-gpt-oss-120b.
