Model registry / Mistral AI / Mistral Small / Mistral Small 4
Mistral Small 4
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Line
Mistral Small
Weights
Open
Released
2026-03-16
Context
256K
Input
$0.15 / 1M
Output
$0.60 / 1M
Max output
256K
Coverage
not covered yet
Overview
Mistral Small 4 is a production model from Mistral AI for chat, extraction, and cost-sensitive agents. It is part of the mistral-small line. It has a context window of 256000 tokens and a max output of 256000 tokens. Weights are open. It supports reasoning mode, tool calling, and accepts image and text inputs, returning text. Price starts at $0.15 per million input tokens and $0.60 per million output tokens, with cached input at $0.015 per million tokens. Knowledge cutoff is 2025-06. It is served by 10 hosts.
Specs
Pricing
Against Mistral Small 3.2: window 128K to 256K, input $0.10 to $0.15, output $0.30 to $0.60.
Range across the hosts that serve it: up to $0.5811 per 1M input tokens.
Strengths
- 256k context and output window
- Reasoning mode and tool calling
- Accepts image and text inputs
- Open weights with low input price
Best for
- Reach for it for cost-sensitive agent workflows
- Reach for it for extracting structured data from text
- Reach for it for chat with long context
- Reach for it for multimodal tasks with image and text
How to access
10 hosts serve this model at the price above · the maker's documentation
Mistral Small: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 256000 tokens, and max output is also 256000 tokens.
- Does it support tool calling?
- Yes, it supports tool calling and reasoning mode.
- What is the price?
- Input from $0.15 per million tokens, output from $0.60 per million tokens, cached input from $0.015 per million tokens.