Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-4.7-Flash
GLM-4.7-Flash
Budget GLM lane for fast coding help, routing, and everyday automation
Line
GLM Flash
Weights
Open
Released
2026-01-19
Context
200K
Input
$0.06 / 1M
Output
$0.40 / 1M
Max output
131K
Coverage
not covered yet
Overview
GLM-4.7-Flash is a budget model in the GLM line for fast coding help, routing, and everyday automation, as described by Z.ai. It sits in the glm-flash line, succeeding GLM-4.6V-Flash. It has a 200,000 token context window and a 131,072 token maximum output. It accepts and returns text. It supports reasoning mode and tool calling. The weights are open. Pricing starts at $0.06 per million input tokens and $0.40 per million output tokens, with cached input from $0.01 per million tokens.
Specs
Pricing
Against GLM-4.6V-Flash: window 128K to 200K, input $0.0218 to $0.06, output $0.2184 to $0.40, cached input $0.0044 to $0.01.
Range across the hosts that serve it: up to $0.08 per 1M input tokens.
Strengths
- 200k token context window
- 131k token max output
- Open weights for self-hosting
- Supports reasoning and tool calling
- Low cost from $0.06 per million input tokens
Best for
- Reach for it for fast coding help
- Reach for it for routing and automation workflows
- Reach for it for budget-friendly text processing
How to access
17 hosts serve this model at the price above · the maker's documentation
GLM Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window of GLM-4.7-Flash?
- It has a 200,000 token context window and a 131,072 token maximum output.
- Is GLM-4.7-Flash open weights?
- Yes, the weights are open. It also supports reasoning mode and tool calling.
- What is the pricing?
- Input from $0.06 per million tokens, output from $0.40 per million tokens, cached input from $0.01 per million tokens.