Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-5.3-Flash
GLM-5.3-Flash
Native multimodal GLM model for efficient coding and long-horizon agent tasks
Line
GLM Flash
Weights
Open
Released
2026-08-26
Context
1M
Input
$0.03 / 1M
Output
$0.025 / 1M
Max output
131K
Coverage
not covered yet
Overview
GLM-5.3-Flash is a native multimodal GLM model from Z.ai (Zhipu, GLM) designed for efficient coding and long-horizon agent tasks. It sits in the glm-flash line, released on 2026-08-26, following the previous version GLM-4.7-FlashX.
Specs
Pricing
Against GLM-4.7-FlashX: window 200K to 1M, input $0.06 to $0.03, output $0.40 to $0.025, cached input $0.01 to $0.006.
Range across the hosts that serve it: up to $0.45 per 1M input tokens.
Strengths
- 1,000,000 token context window
- 131,072 token max output
- Accepts image, pdf, text, and video
- Supports reasoning mode and tool calling
- Open weights with 59 hosting options
Best for
- Reach for it for long-horizon agent tasks
- Reach for it for efficient coding workflows
- Reach for it for multimodal input processing
- Reach for it for low-cost high-volume inference
How to access
59 hosts serve this model at the price above · the maker's documentation
GLM Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and output limit?
- The context window is 1,000,000 tokens and the maximum output is 131,072 tokens.
- Does it support tool calling and reasoning?
- Yes, it supports both reasoning mode and tool calling.
- What is the price per token?
- Input costs from $0.03 per million tokens, output from $0.025 per million, and cached input from $0.006 per million.