25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-5.3-Flash

GLM-5.3-Flash

Native multimodal GLM model for efficient coding and long-horizon agent tasks

Line

GLM Flash

Weights

Open

Released

2026-08-26

Context

1M

Input

$0.03 / 1M

Output

$0.025 / 1M

Max output

131K

Coverage

not covered yet

Overview

GLM-5.3-Flash is a native multimodal GLM model from Z.ai (Zhipu, GLM) designed for efficient coding and long-horizon agent tasks. It sits in the glm-flash line, released on 2026-08-26, following the previous version GLM-4.7-FlashX.

Specs

Released2026-08-26
LineZ.ai (Zhipu, GLM) · GLM Flash
WeightsOpen
Context1M tokens
Max output131K tokens
Inputimage, pdf, text, video
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoffnot stated
API idglm-5.3-flash

Pricing

Input$0.03 / 1M tokens
Cached input$0.006 / 1M tokens
Output$0.025 / 1M tokens

Against GLM-4.7-FlashX: window 200K to 1M, input $0.06 to $0.03, output $0.40 to $0.025, cached input $0.01 to $0.006.

Range across the hosts that serve it: up to $0.45 per 1M input tokens.

Strengths

  • 1,000,000 token context window
  • 131,072 token max output
  • Accepts image, pdf, text, and video
  • Supports reasoning mode and tool calling
  • Open weights with 59 hosting options

Best for

  • Reach for it for long-horizon agent tasks
  • Reach for it for efficient coding workflows
  • Reach for it for multimodal input processing
  • Reach for it for low-cost high-volume inference

How to access

ProviderModel id
302.AIglm-5.3-flash
AIHubMixglm-5.3-flash
Basetenglm-5.3-flash
Berget.AIglm-5.3-flash
Bothubglm-5.3-flash
Charm Hyperglm-5.3-flash
ClinePassglm-5.3-flash
CoreWeaveglm-5.3-flash
Cortecsglm-5.3-flash
CrofAIglm-5.3-flash
CrossModelglm-5.3-flash
Deep Infraglm-5.3-flash

59 hosts serve this model at the price above · the maker's documentation

GLM Flash: every version

VersionReleasedContextInput
GLM-5.3-FlashXCurrent2026-09-181M$0.15
GLM-5.3-Flash2026-08-261M$0.03
GLM-4.7-Flash2026-01-19200K$0.06
GLM-4.7-FlashX2026-01-19200K$0.06
GLM-4.6V-Flash2025-12-08128K$0.0218
GLM-4.5-Flash2025-07-28131K$0.00

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and output limit?
The context window is 1,000,000 tokens and the maximum output is 131,072 tokens.
Does it support tool calling and reasoning?
Yes, it supports both reasoning mode and tool calling.
What is the price per token?
Input costs from $0.03 per million tokens, output from $0.025 per million, and cached input from $0.006 per million.
Open weightsReasoning1M context