GLM 4.7 (Cerebras)
Coding agents that need GLM-class quality with Cerebras-level speed.
What are GLM 4.7 (Cerebras)'s specs and price?
GLM 4.7 (Cerebras), built by Cerebras, ships a 200K-token context window and a 33K-token max output, released 2026-01. It supports text input with a dedicated reasoning mode and costs $2.38 per million blended tokens, the 20th-cheapest of 32 models we track.
Specs
| Context window | 200K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2026-01 |
| Knowledge cutoff | 2025-10 |
| Provider | Cerebras |
Verified 2026-08-08 — source.
Where it ranks
Strengths
- Z.ai GLM 4.7 at wafer-scale speed
- Strong coding performance
- Low latency for agent loops
More about GLM 4.7 (Cerebras)
FAQ
What is GLM 4.7 (Cerebras)'s context window?
GLM 4.7 (Cerebras) has a 200K-token context window and a 33K-token max output — the 21st-largest context of the 32 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-08.
Does GLM 4.7 (Cerebras) support vision or audio input?
No — GLM 4.7 (Cerebras) is text-only as of 2026-08-08.
Does GLM 4.7 (Cerebras) have a reasoning or extended-thinking mode?
Yes — GLM 4.7 (Cerebras) exposes a dedicated reasoning mode for multi-step problems.
When was GLM 4.7 (Cerebras) released, and what is its knowledge cutoff?
GLM 4.7 (Cerebras) was released 2026-01 with a knowledge cutoff of 2025-10.
How much does GLM 4.7 (Cerebras) cost, and who provides it?
GLM 4.7 (Cerebras) is served by Cerebras at $2.38/M blended tokens (3:1 input:output) — the 20th-cheapest of 32 current models. Full pricing breakdown: /llm-api-pricing/cerebras-glm-4-7.
