Model registry / Mistral AI / Voxtral / Voxtral Small
Voxtral Small
Instruct model with native audio input for speech understanding and tool use
Line
Voxtral
Weights
Open
Released
2025-07-15
Context
32K
Input
$0.10 / 1M
Output
$0.30 / 1M
Max output
32K
Coverage
not covered yet
Overview
Voxtral Small is an instruct model from Mistral AI with native audio input for speech understanding and tool use. It is the first version of the voxtral line. It has a context window of 32000 tokens and a maximum output of 32000 tokens. The price starts at $0.10 per million input tokens and $0.30 per million output tokens, with cached input at $0.01 per million tokens. The weights are open. It accepts audio and text and returns text. It does not have a reasoning mode. It supports tool calling. It is served by 4 hosts.
Specs
Pricing
First version of its line in the registry.
Strengths
- Native audio input for speech understanding
- Tool calling for function use
- Open weights for self-hosting
- 32k token context and output
- Accepts audio and text, returns text
Best for
- Reach for it for speech understanding from audio
- Reach for it for tool use in conversations
- Reach for it for building with open weights
- Reach for it for long context tasks up to 32k tokens
How to access
4 hosts serve this model at the price above · the maker's documentation
Voxtral: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and output limit?
- Both are 32000 tokens.
- Does it support tool calling?
- Yes, it supports tool calling.
- Is it open weights?
- Yes, the weights are open.