Model registry / Z.ai (Zhipu, GLM) / GLM / GLM-5-Turbo
GLM-5-Turbo
Faster GLM-5 lane for coding agents that need lower latency
Line
GLM
Weights
API only
Released
2026-03-16
Context
200K
Input
$0.72 / 1M
Output
$3.19 / 1M
Max output
131K
Coverage
not covered yet
Overview
GLM-5-Turbo is a faster lane in the GLM-5 line, designed for coding agents that need lower latency. It is made by Z.ai (Zhipu, GLM) and released on 2026-03-16. It has a context window of 200,000 tokens and a max output of 131,072 tokens. Pricing starts at $0.72 per million input tokens and $3.19 per million output tokens, with cached input at $0.174 per million. Weights are not open. It supports reasoning mode and tool calling, and accepts and returns text. It is served by 18 hosts. Knowledge cutoff is not stated.
Specs
Pricing
Against GLM-5: window 205K to 200K, input $0.50 to $0.72, output $1.92 to $3.19, cached input $0.10 to $0.174.
Range across the hosts that serve it: up to $1.20 per 1M input tokens.
Strengths
- Lower latency for coding agents
- 200,000 token context window
- 131,072 token max output
- Supports reasoning mode
- Supports tool calling
Best for
- Reach for it for coding agent workflows needing low latency
- Reach for it for long-context code analysis
- Reach for it for tool-calling automation
- Reach for it for reasoning-heavy coding tasks
How to access
18 hosts serve this model at the price above · the maker's documentation
GLM: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- How does GLM-5-Turbo compare to GLM-5?
- GLM-5-Turbo has a smaller context window (200,000 vs 204,800) and a higher input price ($0.72 vs $0.50 per million tokens), but it is designed for lower latency.
- Are the weights open?
- No, the weights are not open.
- What is the maximum output length?
- The maximum output is 131,072 tokens.