Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-4.5-Flash
GLM-4.5-Flash
Efficient GLM model for fast reasoning, coding, and agent workflows
Line
GLM Flash
Weights
Open
Released
2025-07-28
Context
131K
Input
$0.00 / 1M
Output
$0.00 / 1M
Max output
98K
Coverage
not covered yet
Overview
GLM-4.5-Flash is an efficient GLM model from Z.ai (Zhipu, GLM) for fast reasoning, coding, and agent workflows. It is the first version of the glm-flash line in the registry, released on 2025-07-28. The model has a context window of 131072 tokens and a maximum output of 98304 tokens. It is priced from $0.00 per million input tokens and $0.00 per million output tokens, with cached input from $0.00 per million tokens. The weights are open. It supports a reasoning mode, tool calling, and accepts and returns text. Its knowledge cutoff is 2025-04, and it is served by 4 hosts.
Specs
Pricing
First version of its line in the registry.
Strengths
- 131072 token context window
- 98304 token max output
- Open weights, reasoning mode, tool calling
- Priced at $0.00 per million tokens
- Text input and output only
Best for
- Reach for it for fast reasoning tasks
- Reach for it for coding workflows
- Reach for it for agent workflows with tool calling
How to access
4 hosts serve this model at the price above · the maker's documentation
GLM Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window size?
- The context window is 131072 tokens, and the maximum output is 98304 tokens.
- Is GLM-4.5-Flash free to use?
- The price is from $0.00 per million input tokens and $0.00 per million output tokens, with cached input from $0.00 per million tokens.
- Does it support tool calling?
- Yes, it supports tool calling and also has a reasoning mode. It accepts and returns text only.