← Back to all pricing

Cerebras GPT-OSS 120B API Pricing: Extreme Speed on Wafer Scale

Analyze Cerebras GPT-OSS 120B API pricing ($0.35/M input, $0.75/M output), extreme wafer-scale inference speed (2,000+ tps), and high-throughput cost efficiency.

Full specs, context window and API limits →

How much does GPT-OSS 120B (Cerebras) cost per million tokens?

Cerebras GPT-OSS 120B costs $0.35 per million input tokens and $0.75 per million output tokens ($0.45/M blended at 3:1). Provides ultra-fast open-weights inference on the Cerebras CS-3 system. Verified 2026-09-08.

Verified 2026-09-08 source
Input
$0.35/M
Output
$0.75/M
Blended
$0.45/M
Provider
Verified 2026-06-14source

How much does GPT-OSS 120B (Cerebras) cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.32× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.1220
Medium1,000500$1.2200
Long4,0002,000$4.8800

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

This owner is the Cerebras-hosted 120B delivery surface. The API bill is joined to the controlled speed sample; quota, hardware, and SLA claims remain unavailable.

1. Fixed agent bill joined to completion time

Agent shape100K billTTFT / throughput sampleEstimated completion time
8,000 input / 1,200 output$370.002450 tokens/sec; TTFT 90 ms; 5 measured samples0.49 sec output-only estimate
32,000 / 4,000 long agent$1420.002450 tokens/sec; TTFT 90 ms; 5 measured samples1.63 sec output-only estimate

Formula: output tokens ÷ measured tokens/sec; this excludes queueing and network time.

2. Request-volume and output-length capacity table

Monthly requestsOutput eachToken billQueue/concurrency boundary
100K500$107.50Concurrency unavailable
1M500$1075.00Concurrency unavailable
1M2,000$2200.00Output length is the sourced sensitivity

3. SDK adoption and quota evidence ledger

EvidenceDated resultSpend calculationMissing boundary
Replay-backed SDK examplesAPI docs / examples$107.50SLA and account quota unavailable
Account quotaUnavailable$107.50Cerebras quota terms not in price record
Self-hostingUnavailableAPI spend is not transferableHardware and utilization unavailable

Verified 2026-06-14. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this Batch 5 scenario.

All three Batch 5 decisions are server-rendered for GPT-OSS 120B (Cerebras); fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.

Batch 68 · exact model pricing decision contributions · verified 2026-09-08

Exact model boundary: Cerebras cerebras-gpt-oss-120b (slug cerebras-gpt-oss-120b). First-party provider pricing and API documentation remain fact owners.

Wafer-scale token pricing and high-throughput monthly spend matrix

Frozen Batch 68 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.35 + out_tokens * 0.75) / 1M) Boundary: Owns Cerebras CS-3 hardware acceleration token tariff modeling.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-gpt-oss-120b-m1-r1
50K instant customer support queries
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=50K instant customer support queries; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50K instant customer support queries is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r2
200K real-time document summarizations
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=200K real-time document summarizations; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 200K real-time document summarizations is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r3
1M high-concurrency event extraction calls
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=1M high-concurrency event extraction calls; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1M high-concurrency event extraction calls is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r4
batch offline ingestion queue
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r5
enterprise dedicated CS-3 provisioning
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=enterprise dedicated CS-3 provisioning; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — enterprise dedicated CS-3 provisioning is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r6
unresolved billing currency
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Ultra-high token throughput SLA and generation turnaround audit

Frozen Batch 68 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Cerebras Wafer-Scale Engine 2,000+ tps Boundary: Owns turnaround time benchmarks and user experience responsiveness.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-gpt-oss-120b-m2-r1
sub-50ms conversational streaming response
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=sub-50ms conversational streaming response; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — sub-50ms conversational streaming response is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r2
rapid code autocomplete (<100ms)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=rapid code autocomplete (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — rapid code autocomplete (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r3
instant multi-page legal summary (<250ms)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=instant multi-page legal summary (<250ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — instant multi-page legal summary (<250ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r4
high-concurrency request surge
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=high-concurrency request surge; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r5
network transit buffer
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=network transit buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — network transit buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r6
unmeasured speed fixture
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Cerebras documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Hardware acceleration cost efficiency and self-hosting break-even

Frozen Batch 68 scenario board. Formula / deterministic rule: managed_cost = tokens * blended_rate; cluster_cost = (8 * H100_hourly + power) * 730 Boundary: Owns cloud wafer-scale API versus private on-prem GPU cluster economics.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-gpt-oss-120b-m3-r1
low intermittent workload (<10M tokens/mo)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=low intermittent workload (<10M tokens/mo); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — low intermittent workload (<10M tokens/mo) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r2
50M monthly tokens (API optimal)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=50M monthly tokens (API optimal); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50M monthly tokens (API optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r3
250M high-throughput enterprise scale
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=250M high-throughput enterprise scale; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 250M high-throughput enterprise scale is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r4
1B+ constant saturation (cluster threshold)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=1B+ constant saturation (cluster threshold); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1B+ constant saturation (cluster threshold) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r5
uncommitted hardware idle hours
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=uncommitted hardware idle hours; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — uncommitted hardware idle hours is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r6
unresolved data center energy tariff
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unresolved data center energy tariff; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved data center energy tariff has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the cerebras-gpt-oss-120b Batch 68 scenario →

Batch 85 Exact-Model Pricing ContributionsOwner: cerebras-gpt-oss-120b (cerebras)verifiedAt=2026-06-14; revalidation required before current claims.

GPT-OSS 120B on Cerebras pricing evidence

Source-backed dated rate shape: dated input=$0.350000/M; output=$0.750000/M. Cerebras-hosted GPT-OSS 120B rate and throughput arithmetic only; Groq hosting, open-weight specifications, and broad speed rankings retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.

Module 1 of 3: Wafer-scale coding bill ladder

Novel contribution boundary: Owns dated Cerebras-hosted 120B arithmetic for fixed coding token shapes; throughput and coding quality are not asserted. Formula / deterministic rule: spend = requests × (inputTokens × inputCostPer1k + outputTokens × outputCostPer1k) / 1,000

Scenario / field IDExact model, provider, dated rates, and fixed inputsResultState
batch85-cerebras-gpt-oss-120b-m1-r1
10K compact coding runs
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=10,000; inputTokens=2,000; outputTokens=500; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch85-cerebras-gpt-oss-120b-m1-r2
2K repository repair runs
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=2,000; inputTokens=20,000; outputTokens=3,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch85-cerebras-gpt-oss-120b-m1-r3
500 long-agent runs
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=500; inputTokens=80,000; outputTokens=8,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://www.cerebras.ai/pricing; registry owner=cerebras; verifiedAt=2026-06-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 2 of 3: Cerebras→Groq accepted-run cost threshold

Novel contribution boundary: Owns dated hosted-cost arithmetic with explicit retry input; accepted-run quality, throughput, and provider behavior are not sourced. Formula / deterministic rule: requiredAcceptedRate = GroqAttemptCost / CerebrasAttemptCost; observed acceptance = Unavailable

Scenario / field IDExact model, provider, dated rates, and fixed inputsResultState
batch85-cerebras-gpt-oss-120b-m2-r1
2K-input coding run; 10% retry share; 500 retry-input tokens
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1; inputTokens=2,000; outputTokens=500; retryInputTokens=500; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch85-cerebras-gpt-oss-120b-m2-r2
20K-input repair run; 20% retry share; 2,000 retry-input tokens
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1; inputTokens=20,000; outputTokens=3,000; retryInputTokens=2,000; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch85-cerebras-gpt-oss-120b-m2-r3
80K-input agent run; 30% retry share; 4,000 retry-input tokens
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1; inputTokens=80,000; outputTokens=8,000; retryInputTokens=4,000; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://www.cerebras.ai/pricing; registry owner=cerebras; verifiedAt=2026-06-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 3 of 3: Throughput, capacity, and deployment evidence ledger

Novel contribution boundary: Owns exact-model Cerebras pricing evidence only; wafer-scale capacity, quotas, deployment details, and SLAs are not fabricated. Formula / deterministic rule: evidence = exact-model registry field when present; throughput, quota, capacity, and SLA inputs = Unavailable

Scenario / field IDExact model, provider, dated rates, and fixed inputsResultState
batch85-cerebras-gpt-oss-120b-m3-r1
Exact-model throughput field
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=evidenceField=throughput; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch85-cerebras-gpt-oss-120b-m3-r2
Account quota or capacity field
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=evidenceField=quota; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
batch85-cerebras-gpt-oss-120b-m3-r3
Deployment or SLA term
model=cerebras-gpt-oss-120b; provider=cerebras; registryRates=dated input=$0.350000/M; output=$0.750000/M; comparisonModel=gpt-oss-120b; fixedInputs=evidenceField=SLA; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — verifiedAt=2026-06-14 is before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://www.cerebras.ai/pricing; registry owner=cerebras; verifiedAt=2026-06-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Try GPT-OSS 120B on Cerebras pricing analysis →

How fast is GPT-OSS 120B (Cerebras)?

Tokens / sec
2450
TTFT
90 ms
Rank
#1 of 31
$ / M ÷ t/s
$0.0002
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does GPT-OSS 120B (Cerebras) cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.04
1,000,000$0.45
10,000,000$4.50
100,000,000$45.00

How does GPT-OSS 120B (Cerebras) compare with other models?

GLM 4.7 (Cerebras)$2.38/MCodestral$0.45/MGPT-5.4 Nano$0.46/MGemini 3.1 Flash Lite$0.56/M
See all Cerebras models →

What is GPT-OSS 120B (Cerebras) best for?

#3 for Writing & Content#3 for Agents & Tool Use#3 for Chatbots & Support
Looking for a cheaper option?
GPT-OSS 20B is 70.8% cheaper — a config migration. See all 8 alternatives to GPT-OSS 120B (Cerebras)

What are common questions about GPT-OSS 120B (Cerebras)?

Is GPT-OSS 120B (Cerebras) cheaper than Codestral?

GPT-OSS 120B (Cerebras) costs $0.45/M blended tokens, Codestral costs $0.45/M — Codestral is cheaper.

How much does 1 million tokens cost with GPT-OSS 120B (Cerebras)?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.45. Pure input costs $0.35/M; pure output costs $0.75/M.

What does GPT-OSS 120B (Cerebras) cost at high volume?

At 100 million blended tokens a month, GPT-OSS 120B (Cerebras) costs approximately $45.00. See the cost-at-scale table below for other volumes.

Try GPT-OSS 120B (Cerebras) for free

Run real prompts against GPT-OSS 120B (Cerebras) and every other model on this page in one workspace.

Try GPT-OSS 120B (Cerebras) Free