Model registry / Z.ai (Zhipu, GLM) / GLM Flash / GLM-4.6V-Flash
GLM-4.6V-Flash
Lightweight GLM vision model for visual reasoning, documents, and multimodal agents
Line
GLM Flash
Weights
Open
Released
2025-12-08
Context
128K
Input
$0.0218 / 1M
Output
$0.2184 / 1M
Max output
33K
Coverage
not covered yet
Overview
GLM-4.6V-Flash is a lightweight vision model from Z.ai for visual reasoning, documents, and multimodal agents. It is part of the glm-flash line, succeeding GLM-4.5-Flash. It has a 128000-token context window and a 32768-token output limit. Weights are open. It supports reasoning mode and tool calling. It accepts image, text, and video inputs and returns text. Pricing starts at $0.0218 per million input tokens and $0.2184 per million output tokens, with cached input at $0.0044 per million tokens. It is served by 5 hosts.
Specs
Pricing
Against GLM-4.5-Flash: window 131K to 128K, input $0.00 to $0.0218, output $0.00 to $0.2184, cached input $0.00 to $0.0044.
Range across the hosts that serve it: up to $0.30 per 1M input tokens.
Strengths
- Visual reasoning on images, text, and video
- Document understanding for multimodal tasks
- Supports reasoning mode and tool calling
- Open weights for self-hosting
- 128k context window with 32k output
Best for
- Reach for it for visual reasoning on images and video
- Reach for it for document analysis and extraction
- Reach for it for building multimodal agents
- Reach for it for tasks needing tool calling and reasoning
How to access
5 hosts serve this model at the price above · the maker's documentation
GLM Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and output limit?
- The context window is 128000 tokens and the maximum output is 32768 tokens.
- Does it support video input?
- Yes, it accepts image, text, and video inputs, and returns text.
- Is the model open weights?
- Yes, the weights are open. It also supports reasoning mode and tool calling.