Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-4.7-FlashX
GLM-4.7-FlashX
Efficient GLM model for fast reasoning, coding, and agent workflows
Line
GLM Flash
Weights
Open
Released
2026-01-19
Context
200K
Input
$0.06 / 1M
Output
$0.40 / 1M
Max output
131K
Coverage
not covered yet
Overview
GLM-4.7-FlashX is an efficient GLM model for fast reasoning, coding, and agent workflows. It belongs to the glm-flash line, succeeding GLM-4.7-Flash. It has a context window of 200,000 tokens and a maximum output of 131,072 tokens. It accepts and returns text. It supports reasoning mode and tool calling. The weights are open. Pricing starts at $0.06 per million input tokens and $0.40 per million output tokens, with cached input at $0.01 per million tokens. It is served by 9 hosts.
Specs
Pricing
Against GLM-4.6V-Flash: window 128K to 200K, input $0.0218 to $0.06, output $0.2184 to $0.40, cached input $0.0044 to $0.01.
Range across the hosts that serve it: up to $0.0728 per 1M input tokens.
Strengths
- 200k token context window
- 131k token max output
- Open weights for self-hosting
- Supports reasoning and tool calling
Best for
- Reach for it for fast reasoning tasks
- Reach for it for coding assistance
- Reach for it for agent workflows
How to access
9 hosts serve this model at the price above · the maker's documentation
GLM Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 200,000 tokens.
- Does it support tool calling?
- Yes, it supports tool calling and reasoning mode.
- Is the model open source?
- Yes, the weights are open.