Model registry / Xiaomi / MiMo / MiMo-V2.6-Flash
MiMo-V2.6-Flash
MiMo Flash model for multimodal coding agents and long-context automation
Line
MiMo
Weights
Open
Released
2026-09-22
Context
1.05M
Input
$0.04 / 1M
Output
$0.28 / 1M
Max output
131K
Coverage
not covered yet
Overview
MiMo-V2.6-Flash is a multimodal model from Xiaomi for coding agents and long-context automation. It is part of the MiMo line, succeeding MiMo-V2.5-Pro. It has a context window of 1,048,576 tokens and a maximum output of 131,072 tokens. It accepts audio, image, text, and video, and returns text. It supports reasoning mode and tool calling. Weights are open. Pricing starts at $0.04 per million input tokens and $0.28 per million output tokens, with cached input at $0.0027 per million tokens.
Specs
Pricing
Against MiMo-V2.5-Pro: input $0.40 to $0.04, output $0.80 to $0.28, cached input $0.003 to $0.0027.
Range across the hosts that serve it: up to $0.1692 per 1M input tokens.
Strengths
- 1M token context window for long documents
- Accepts audio, image, text, and video inputs
- Supports reasoning mode and tool calling
- Open weights for customization
- 131k token max output
Best for
- Reach for it for coding agents that need long context
- Reach for it for multimodal automation pipelines
- Reach for it for processing long videos or documents
- Reach for it for tool-calling workflows
How to access
19 hosts serve this model at the price above · the maker's documentation
MiMo: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window size?
- The context window is 1,048,576 tokens, and the maximum output is 131,072 tokens.
- Does it support image and video input?
- Yes, it accepts audio, image, text, and video inputs, and returns text output.
- Are the weights open?
- Yes, the weights are open. It also supports reasoning mode and tool calling.