Model registry / Alibaba (Qwen) / Qwen / Qwen3.8 Flash
Qwen3.8 Flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
Line
Qwen
Weights
API only
Released
2026-08-26
Context
1M
Input
$0.11 / 1M
Output
$0.38 / 1M
Max output
131K
Coverage
not covered yet
Overview
Qwen3.8 Flash is a vision-language model from Alibaba's Qwen line, designed for visual reasoning, documents, and agent tasks. It is a later release in the Qwen line, following Qwen3.8 Max. It has a context window of 1,000,000 tokens and a maximum output of 131,072 tokens. Pricing starts at $0.11 per million input tokens and $0.38 per million output tokens, with cached input at $0.011 per million tokens. The weights are not open. It supports reasoning mode, tool calling, and accepts image, text, and video input, returning text output.
Specs
Pricing
Against Qwen3.8 Max: input $0.338 to $0.11, output $1.01 to $0.38, cached input $0.0676 to $0.011.
Range across the hosts that serve it: up to $0.18 per 1M input tokens.
Strengths
- 1M token context window
- 131K max output tokens
- Accepts image, text, and video
- Reasoning mode and tool calling
- Priced from $0.11/M input tokens
Best for
- Reach for it for visual reasoning tasks
- Reach for it for document analysis
- Reach for it for agent tasks with tool calling
How to access
25 hosts serve this model at the price above · the maker's documentation
Qwen: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 1,000,000 tokens.
- Does it support tool calling?
- Yes, it supports tool calling and reasoning mode.
- Is it open weights?
- No, the weights are not open.