Model registry / Inception / Mercury / Mercury 2.5
Mercury 2.5
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception
Line
Mercury
Weights
API only
Released
2026-09-08
Context
260K
Input
$0.04 / 1M
Output
$0.15 / 1M
Max output
66K
Coverage
not covered yet
Overview
Mercury 2.5 is a reasoning LLM and diffusion LLM (dLLM) from Inception. It is the latest model in the mercury line and is marketed as the fastest reasoning LLM. It has a context window of 260,000 tokens and a maximum output of 65,536 tokens. Pricing starts at $0.04 per million input tokens and $0.15 per million output tokens, with cached input at $0.004 per million. The weights are not open. It supports reasoning mode and tool calling, and accepts and returns text.
Specs
Pricing
Against Mercury Edit 2: window 32K to 260K, input $0.25 to $0.04, output $0.75 to $0.15, cached input $0.025 to $0.004.
Range across the hosts that serve it: up to $0.20 per 1M input tokens.
Strengths
- 260k token context window
- 65k token max output
- Reasoning mode and tool calling
- Input from $0.04 per million tokens
- Cached input at $0.004 per million
Best for
- Reach for it for long-context reasoning tasks
- Reach for it for tool-calling workflows
- Reach for it for cost-efficient text generation
- Reach for it for large output generation
How to access
5 hosts serve this model at the price above · the maker's documentation
Mercury: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window of Mercury 2.5?
- 260,000 tokens.
- Does Mercury 2.5 support tool calling?
- Yes, it supports tool calling and reasoning mode.
- What is the pricing for Mercury 2.5?
- Input from $0.04 per million tokens, output from $0.15 per million, cached input at $0.004 per million.