Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-5.3-FlashX
GLM-5.3-FlashX
High-speed GLM-5.3-Flash serving option for coding and agent workflows
Line
GLM Flash
Weights
Open
Released
2026-09-18
Context
1M
Input
$0.15 / 1M
Output
$0.50 / 1M
Max output
131K
Coverage
not covered yet
Overview
GLM-5.3-FlashX is a high-speed serving option from Z.ai for coding and agent workflows. It is part of the glm-flash line, released after GLM-5.3-Flash on 2026-09-18. It has a 1,000,000 token context window and a 131,072 token output limit. Weights are open. It supports reasoning mode and tool calling. It accepts image, pdf, text, and video, and returns text. Price starts at $0.15 per million input tokens and $0.50 per million output tokens, with cached input at $0.03 per million tokens.
Specs
Pricing
Against GLM-5.3-Flash: input $0.03 to $0.15, output $0.025 to $0.50, cached input $0.006 to $0.03.
Range across the hosts that serve it: up to $0.375 per 1M input tokens.
Strengths
- 1M token context window
- 131K token max output
- Open weights
- Supports reasoning and tool calling
- Accepts image, pdf, text, video
Best for
- Reach for it for coding and agent workflows
- Reach for it for long-context tasks
- Reach for it for multimodal input processing
How to access
8 hosts serve this model at the price above · the maker's documentation
GLM Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window size?
- The context window is 1,000,000 tokens.
- Is GLM-5.3-FlashX open weights?
- Yes, the weights are open.
- What is the pricing?
- From $0.15 per million input tokens and $0.50 per million output tokens. Cached input is $0.03 per million tokens.