Model registry / NVIDIA / Nemotron / Nemotron 3 Super
Nemotron 3 Super
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
Line
Nemotron
Weights
Open
Released
2026-03-11
Context
262K
Input
$0.05 / 1M
Output
$0.25 / 1M
Max output
262K
Coverage
not covered yet
Overview
Nemotron 3 Super is the middle tier in NVIDIA's Nemotron line, designed for collaborative agents and high-volume reasoning workloads. It was released on 2026-03-11. It has a context window of 262,144 tokens and a maximum output of 262,144 tokens. Weights are open. It supports reasoning mode and tool calling. It accepts text and returns text. Pricing starts at $0.05 per million input tokens and $0.25 per million output tokens, with cached input at $0.025 per million tokens. It is served by 9 hosts.
Specs
Pricing
Against Nemotron Nano 12B v2 VL: window 128K to 262K, input $0.20 to $0.05, output $0.60 to $0.25, cached input $0.04 to $0.025.
Range across the hosts that serve it: up to $0.30 per 1M input tokens.
Strengths
- 262k token context window
- 262k token max output
- Supports reasoning mode
- Supports tool calling
- Open weights for customization
Best for
- Reach for it for collaborative agent workflows
- Reach for it for high-volume reasoning tasks
- Reach for it for long-context text generation
- Reach for it for tool-augmented responses
How to access
9 hosts serve this model at the price above · the maker's documentation
Nemotron: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and output limit?
- The context window is 262,144 tokens and the maximum output is also 262,144 tokens.
- What is the pricing?
- Input starts at $0.05 per million tokens, output at $0.25 per million, and cached input at $0.025 per million.
- Does it support reasoning and tool calling?
- Yes, it has a reasoning mode and supports tool calling.