25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-4.6V-Flash

GLM-4.6V-Flash

Lightweight GLM vision model for visual reasoning, documents, and multimodal agents

Line

GLM Flash

Weights

Open

Released

2025-12-08

Context

128K

Input

$0.0218 / 1M

Output

$0.2184 / 1M

Max output

33K

Coverage

not covered yet

Overview

GLM-4.6V-Flash is a lightweight vision model from Z.ai for visual reasoning, documents, and multimodal agents. It is part of the glm-flash line, succeeding GLM-4.5-Flash. It has a 128000-token context window and a 32768-token output limit. Weights are open. It supports reasoning mode and tool calling. It accepts image, text, and video inputs and returns text. Pricing starts at $0.0218 per million input tokens and $0.2184 per million output tokens, with cached input at $0.0044 per million tokens. It is served by 5 hosts.

Specs

Released2025-12-08
LineZ.ai (Zhipu, GLM) · GLM Flash
WeightsOpen
Context128K tokens
Max output33K tokens
Inputimage, text, video
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoffnot stated
API idglm-4.6v-flash

Pricing

Input$0.0218 / 1M tokens
Cached input$0.0044 / 1M tokens
Output$0.2184 / 1M tokens

Against GLM-4.5-Flash: window 131K to 128K, input $0.00 to $0.0218, output $0.00 to $0.2184, cached input $0.00 to $0.0044.

Range across the hosts that serve it: up to $0.30 per 1M input tokens.

Strengths

  • Visual reasoning on images, text, and video
  • Document understanding for multimodal tasks
  • Supports reasoning mode and tool calling
  • Open weights for self-hosting
  • 128k context window with 32k output

Best for

  • Reach for it for visual reasoning on images and video
  • Reach for it for document analysis and extraction
  • Reach for it for building multimodal agents
  • Reach for it for tasks needing tool calling and reasoning

How to access

ProviderModel id
Hugging Faceglm-4.6v-flash
Tempr Gatewayglm-4.6v-flash
Z.AIglm-4.6v-flash
ZenMuxglm-4.6v-flash
Zhipu AIglm-4.6v-flash

5 hosts serve this model at the price above · the maker's documentation

GLM Flash: every version

VersionReleasedContextInput
GLM-5.3-FlashXCurrent2026-09-181M$0.15
GLM-5.3-Flash2026-08-261M$0.03
GLM-4.7-Flash2026-01-19200K$0.06
GLM-4.7-FlashX2026-01-19200K$0.06
GLM-4.6V-Flash2025-12-08128K$0.0218
GLM-4.5-Flash2025-07-28131K$0.00

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and output limit?
The context window is 128000 tokens and the maximum output is 32768 tokens.
Does it support video input?
Yes, it accepts image, text, and video inputs, and returns text.
Is the model open weights?
Yes, the weights are open. It also supports reasoning mode and tool calling.
Open weightsReasoning128K context