25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Z.ai (Zhipu, GLM) / GLM / GLM-5-Turbo

GLM-5-Turbo

Faster GLM-5 lane for coding agents that need lower latency

Line

GLM

Weights

API only

Released

2026-03-16

Context

200K

Input

$0.72 / 1M

Output

$3.19 / 1M

Max output

131K

Coverage

not covered yet

Overview

GLM-5-Turbo is a faster lane in the GLM-5 line, designed for coding agents that need lower latency. It is made by Z.ai (Zhipu, GLM) and released on 2026-03-16. It has a context window of 200,000 tokens and a max output of 131,072 tokens. Pricing starts at $0.72 per million input tokens and $3.19 per million output tokens, with cached input at $0.174 per million. Weights are not open. It supports reasoning mode and tool calling, and accepts and returns text. It is served by 18 hosts. Knowledge cutoff is not stated.

Specs

Released2026-03-16
LineZ.ai (Zhipu, GLM) · GLM
WeightsAPI only
Context200K tokens
Max output131K tokens
Inputtext
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoffnot stated
API idglm-5-turbo

Pricing

Input$0.72 / 1M tokens
Cached input$0.174 / 1M tokens
Output$3.19 / 1M tokens

Against GLM-5: window 205K to 200K, input $0.50 to $0.72, output $1.92 to $3.19, cached input $0.10 to $0.174.

Range across the hosts that serve it: up to $1.20 per 1M input tokens.

Strengths

  • Lower latency for coding agents
  • 200,000 token context window
  • 131,072 token max output
  • Supports reasoning mode
  • Supports tool calling

Best for

  • Reach for it for coding agent workflows needing low latency
  • Reach for it for long-context code analysis
  • Reach for it for tool-calling automation
  • Reach for it for reasoning-heavy coding tasks

How to access

ProviderModel id
302.AIglm-5-turbo
Cortecsglm-5-turbo
CrossModelglm-5-turbo
DevPass (LLM Gateway)glm-5-turbo
Eden AIglm-5-turbo
Impossiblglm-5-turbo
Kilo Gatewayglm-5-turbo
LLM Gatewayglm-5-turbo
Merge Gatewayglm-5-turbo
NanoGPTglm-5-turbo
Ofoxglm-5-turbo
OpenRouterglm-5-turbo

18 hosts serve this model at the price above · the maker's documentation

GLM: every version

VersionReleasedContextInput
GLM-5.3Current2026-08-141M$0.40
GLM-5.22026-06-131M$0.30
GLM-5.12026-04-07200K$0.45
GLM-5V-Turbo2026-04-01200K$0.7042
GLM-5-Turbo2026-03-16200K$0.72
GLM-52026-02-12205K$0.50

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

How does GLM-5-Turbo compare to GLM-5?
GLM-5-Turbo has a smaller context window (200,000 vs 204,800) and a higher input price ($0.72 vs $0.50 per million tokens), but it is designed for lower latency.
Are the weights open?
No, the weights are not open.
What is the maximum output length?
The maximum output is 131,072 tokens.
ProprietaryReasoning200K context