Model registry / Mistral AI / Mistral Small / Mistral Small
Mistral Small
Efficient Mistral model for fast chat, extraction, and production assistants
Line
Mistral Small
Weights
Open
Released
2026-03-16
Context
256K
Input
$0.075 / 1M
Output
$0.20 / 1M
Max output
256K
Coverage
not covered yet
Overview
Mistral Small is an efficient model from Mistral AI for fast chat, extraction, and production assistants. It is part of the mistral-small line, succeeding Mistral Small 3.2. It has a context window of 256000 tokens and a max output of 256000 tokens. Weights are open. It supports reasoning mode, tool calling, and accepts image and text inputs, returning text. Pricing starts at $0.075 per million input tokens and $0.20 per million output tokens, with cached input at $0.015 per million tokens. Knowledge cutoff is 2025-06.
Specs
Pricing
Against Mistral Small 3.2: window 128K to 256K, input $0.10 to $0.075, output $0.30 to $0.20.
Range across the hosts that serve it: up to $0.15 per 1M input tokens.
Strengths
- 256k token context and output window
- Open weights for self-hosting
- Supports reasoning and tool calling
- Accepts image and text inputs
- Priced from $0.075 per million input tokens
Best for
- Reach for it for fast chat assistants
- Reach for it for extraction from text and images
- Reach for it for production assistants with tool use
- Reach for it for long-context reasoning tasks
How to access
6 hosts serve this model at the price above · the maker's documentation
Mistral Small: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window size?
- The context window is 256000 tokens, and the max output is also 256000 tokens.
- Does it support image input?
- Yes, it accepts image and text inputs, and returns text output.
- What is the pricing?
- Input starts at $0.075 per million tokens, output at $0.20 per million tokens, and cached input at $0.015 per million tokens.