25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Mistral AI / Mistral Small / Mistral Small 3.2

Mistral Small 3.2

Efficient Mistral model for fast chat, extraction, and production assistants

Line

Mistral Small

Weights

Open

Released

2025-06-20

Context

128K

Input

$0.10 / 1M

Output

$0.30 / 1M

Max output

16K

Coverage

not covered yet

Overview

Mistral Small 3.2 is an efficient Mistral model for fast chat, extraction, and production assistants. It is the first version of the mistral-small line in the registry. It has a context window of 128,000 tokens and a maximum output of 16,384 tokens. Pricing starts at $0.10 per million input tokens and $0.30 per million output tokens. The weights are open. It accepts image and text inputs and returns text. It supports tool calling but not reasoning mode.

Specs

Released2025-06-20
LineMistral AI · Mistral Small
WeightsOpen
Context128K tokens
Max output16K tokens
Inputimage, text
Outputtext
ReasoningNo
Tool callingYes
Knowledge cutoff2025-03
API idmistral-small-2506

Pricing

Input$0.10 / 1M tokens
Cached inputnot stated
Output$0.30 / 1M tokens

First version of its line in the registry.

Strengths

  • 128k context window
  • 16k max output tokens
  • Tool calling supported
  • Accepts image and text inputs
  • Open weights

Best for

  • Reach for it for fast chat assistants
  • Reach for it for extraction from text and images
  • Reach for it for production assistant workflows
  • Reach for it for tool-calling applications

How to access

ProviderModel id
DevPass (LLM Gateway)mistral-small-2506
LLM Gatewaymistral-small-2506
Mistralmistral-small-2506

3 hosts serve this model at the price above · the maker's documentation

Mistral Small: every version

VersionReleasedContextInput
Mistral SmallCurrent2026-03-16256K$0.075
Mistral Small 42026-03-16256K$0.15
Mistral Small 3.22025-06-20128K$0.10

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and output limit?
The context window is 128,000 tokens and the maximum output is 16,384 tokens.
Does it support reasoning mode?
No, reasoning mode is not supported. It does support tool calling.
What is the price per token?
Pricing starts at $0.10 per million input tokens and $0.30 per million output tokens.
Open weights128K context