25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / NVIDIA / Nemotron / Nemotron 3 Super

Nemotron 3 Super

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

Line

Nemotron

Weights

Open

Released

2026-03-11

Context

262K

Input

$0.05 / 1M

Output

$0.25 / 1M

Max output

262K

Coverage

not covered yet

Overview

Nemotron 3 Super is the middle tier in NVIDIA's Nemotron line, designed for collaborative agents and high-volume reasoning workloads. It was released on 2026-03-11. It has a context window of 262,144 tokens and a maximum output of 262,144 tokens. Weights are open. It supports reasoning mode and tool calling. It accepts text and returns text. Pricing starts at $0.05 per million input tokens and $0.25 per million output tokens, with cached input at $0.025 per million tokens. It is served by 9 hosts.

Specs

Released2026-03-11
LineNVIDIA · Nemotron
WeightsOpen
Context262K tokens
Max output262K tokens
Inputtext
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoff2024-04
API idnvidia/nemotron-3-super-120b-a12b

Pricing

Input$0.05 / 1M tokens
Cached input$0.025 / 1M tokens
Output$0.25 / 1M tokens

Against Nemotron Nano 12B v2 VL: window 128K to 262K, input $0.20 to $0.05, output $0.60 to $0.25, cached input $0.04 to $0.025.

Range across the hosts that serve it: up to $0.30 per 1M input tokens.

Strengths

  • 262k token context window
  • 262k token max output
  • Supports reasoning mode
  • Supports tool calling
  • Open weights for customization

Best for

  • Reach for it for collaborative agent workflows
  • Reach for it for high-volume reasoning tasks
  • Reach for it for long-context text generation
  • Reach for it for tool-augmented responses

How to access

ProviderModel id
Kenarinvidia/nemotron-3-super-120b-a12b
Kilo Gatewaynvidia/nemotron-3-super-120b-a12b
NanoGPTnvidia/nemotron-3-super-120b-a12b
Nebius Token Factorynvidia/nemotron-3-super-120b-a12b
Nvidianvidia/nemotron-3-super-120b-a12b
OpenRouternvidia/nemotron-3-super-120b-a12b
Perplexity Agentnvidia/nemotron-3-super-120b-a12b
Requestynvidia/nemotron-3-super-120b-a12b
Vercel AI Gatewaynvidia/nemotron-3-super-120b-a12b

9 hosts serve this model at the price above · the maker's documentation

Nemotron: every version

VersionReleasedContextInput
Nemotron 3 Nano Omni2026-04-28256K$0.10
Nemotron 3 Super2026-03-11262K$0.05
Nemotron Nano 12B v2 VL2025-10-28128K$0.20
nvidia-nemotron-nano-9b-v22025-08-18131K$0.00

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and output limit?
The context window is 262,144 tokens and the maximum output is also 262,144 tokens.
What is the pricing?
Input starts at $0.05 per million tokens, output at $0.25 per million, and cached input at $0.025 per million.
Does it support reasoning and tool calling?
Yes, it has a reasoning mode and supports tool calling.
Open weightsReasoning262K context