25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-5.3-FlashX

GLM-5.3-FlashX

High-speed GLM-5.3-Flash serving option for coding and agent workflows

Line

GLM Flash

Weights

Open

Released

2026-09-18

Context

1M

Input

$0.15 / 1M

Output

$0.50 / 1M

Max output

131K

Coverage

not covered yet

Overview

GLM-5.3-FlashX is a high-speed serving option from Z.ai for coding and agent workflows. It is part of the glm-flash line, released after GLM-5.3-Flash on 2026-09-18. It has a 1,000,000 token context window and a 131,072 token output limit. Weights are open. It supports reasoning mode and tool calling. It accepts image, pdf, text, and video, and returns text. Price starts at $0.15 per million input tokens and $0.50 per million output tokens, with cached input at $0.03 per million tokens.

Specs

Released2026-09-18
LineZ.ai (Zhipu, GLM) · GLM Flash
WeightsOpen
Context1M tokens
Max output131K tokens
Inputimage, pdf, text, video
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoffnot stated
API idglm-5.3-flashx

Pricing

Input$0.15 / 1M tokens
Cached input$0.03 / 1M tokens
Output$0.50 / 1M tokens

Against GLM-5.3-Flash: input $0.03 to $0.15, output $0.025 to $0.50, cached input $0.006 to $0.03.

Range across the hosts that serve it: up to $0.375 per 1M input tokens.

Strengths

  • 1M token context window
  • 131K token max output
  • Open weights
  • Supports reasoning and tool calling
  • Accepts image, pdf, text, video

Best for

  • Reach for it for coding and agent workflows
  • Reach for it for long-context tasks
  • Reach for it for multimodal input processing

How to access

ProviderModel id
Kilo Gatewayglm-5.3-flashx
Ofoxglm-5.3-flashx
OpenRouterglm-5.3-flashx
Tempr Gatewayglm-5.3-flashx
Vercel AI Gatewayglm-5.3-flashx
Z.AIglm-5.3-flashx
ZenMuxglm-5.3-flashx
Zhipu AIglm-5.3-flashx

8 hosts serve this model at the price above · the maker's documentation

GLM Flash: every version

VersionReleasedContextInput
GLM-5.3-FlashXCurrent2026-09-181M$0.15
GLM-5.3-Flash2026-08-261M$0.03
GLM-4.7-Flash2026-01-19200K$0.06
GLM-4.7-FlashX2026-01-19200K$0.06
GLM-4.6V-Flash2025-12-08128K$0.0218
GLM-4.5-Flash2025-07-28131K$0.00

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window size?
The context window is 1,000,000 tokens.
Is GLM-5.3-FlashX open weights?
Yes, the weights are open.
What is the pricing?
From $0.15 per million input tokens and $0.50 per million output tokens. Cached input is $0.03 per million tokens.
Open weightsReasoning1M context