Model registry / NVIDIA / Nemotron / Nemotron 3 Ultra 550B A55B
Nemotron 3 Ultra 550B A55B
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Line
Nemotron
Weights
Open
Released
2026-06-04
Context
1M
Input
$0.50 / 1M
Output
$2.20 / 1M
Max output
66K
Coverage
not covered yet
Overview
Nemotron 3 Ultra 550B A55B is the largest model in the Nemotron 3 line, designed for maximum open-weight reasoning and agent accuracy. It is the successor to Nemotron 3 Nano Omni, released earlier in the same line. It has a context window of 1,000,000 tokens and a maximum output of 65,536 tokens. Weights are open. It supports reasoning mode and tool calling, and accepts and returns text. Pricing starts at $0.50 per million input tokens and $2.20 per million output tokens, with cached input at $0.10 per million tokens.
Specs
Pricing
Against Nemotron 3 Nano Omni: window 256K to 1M, input $0.10 to $0.50, output $0.25 to $2.20.
Range across the hosts that serve it: up to $1.00 per 1M input tokens.
Strengths
- 1 million token context window
- 65,536 token max output
- Open weights for self-hosting
- Reasoning mode for complex tasks
- Tool calling for agent workflows
Best for
- Reach for it for complex reasoning with very long context
- Reach for it for building agents that call tools
- Reach for it for open-weight deployment with large output needs
How to access
12 hosts serve this model at the price above · the maker's documentation
Nemotron: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 1,000,000 tokens.
- Does it support tool calling?
- Yes, it supports tool calling.
- What is the pricing?
- From $0.50 per million input tokens and $2.20 per million output tokens, with cached input at $0.10 per million tokens.