GLM-5.3 Flash
Coding agents and visual feedback loops where Flash economics matter.
What are GLM-5.3 Flash's specs and price?
GLM-5.3 Flash, built by Z.ai, ships a 1M-token context window and a 131K-token max output, released 2026-08. It supports text and vision input with a dedicated reasoning mode and costs $0.24 per million blended tokens, the 9th-cheapest of 49 models we track.
What are GLM-5.3 Flash's specs?
| Context window | 1M tokens |
| Max output | 131K tokens |
| Modalities | text, vision |
| Extended thinking | Yes |
| Released | 2026-08 |
| Knowledge cutoff | Not published |
| Provider | Z.ai |
Verified 2026-10-01 — source.
Where does GLM-5.3 Flash rank?
What are GLM-5.3 Flash's strengths?
- First natively multimodal GLM-5 model
- Beats GLM-5.2 at one-tenth the price
- 320B MoE with linear+sparse attention for 1M context
What else should you know about GLM-5.3 Flash?
What are common questions about GLM-5.3 Flash?
What is GLM-5.3 Flash's context window?
GLM-5.3 Flash has a 1M-token context window and a 131K-token max output — the 22nd-largest context of the 49 current models we track. Source: https://docs.z.ai/guides/llm/glm-5.3-flash, verified 2026-10-01.
Does GLM-5.3 Flash support vision or audio input?
Yes — GLM-5.3 Flash accepts vision input in addition to text.
Does GLM-5.3 Flash have a reasoning or extended-thinking mode?
Yes — GLM-5.3 Flash exposes a dedicated reasoning mode for multi-step problems.
When was GLM-5.3 Flash released, and what is its knowledge cutoff?
GLM-5.3 Flash was released 2026-08.
How much does GLM-5.3 Flash cost, and who provides it?
GLM-5.3 Flash is served by Z.ai at $0.24/M blended tokens (3:1 input:output) — the 9th-cheapest of 49 current models. Full pricing breakdown: /llm-api-pricing/glm-5-3-flash.
