Model registry / NVIDIA / Nemotron / Nemotron Nano 12B v2 VL
Nemotron Nano 12B v2 VL
Nemotron multimodal model for visual reasoning and agentic AI workflows
Line
Nemotron
Weights
Open
Released
2025-10-28
Context
128K
Input
$0.20 / 1M
Output
$0.60 / 1M
Max output
128K
Coverage
not covered yet
Overview
Nemotron Nano 12B v2 VL is a multimodal model from NVIDIA for visual reasoning and agentic AI workflows. It is part of the nemotron line. It has a context window of 128,000 tokens and a maximum output of 128,000 tokens. It accepts image, text, and video inputs and returns text. It supports reasoning mode and tool calling. The weights are open. Pricing starts at $0.20 per million input tokens and $0.60 per million output tokens, with cached input at $0.04 per million tokens.
Specs
Pricing
Against nvidia-nemotron-nano-9b-v2: window 131K to 128K, input $0.00 to $0.20, output $0.00 to $0.60.
Strengths
- Accepts image, text, and video inputs
- Supports reasoning mode for complex tasks
- Supports tool calling for agentic workflows
- 128k token context window
- Open weights for customization
Best for
- Reach for it for visual reasoning tasks
- Reach for it for building agentic AI workflows
- Reach for it for multimodal understanding with large context
- Reach for it for tool-calling applications
How to access
3 hosts serve this model at the price above · the maker's documentation
Nemotron: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What input types does it support?
- It accepts image, text, and video inputs and returns text.
- Does it support tool calling?
- Yes, it supports tool calling.
- What is the pricing?
- Input from $0.20 per million tokens, output from $0.60 per million tokens, cached input from $0.04 per million tokens.