Model registry / Google / Gemma / Gemma 4 31B IT
Gemma 4 31B IT
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
Line
Gemma
Weights
Open
Released
2026-04-02
Context
262K
Input
$0.09 / 1M
Output
$0.25 / 1M
Max output
33K
Coverage
not covered yet
Overview
Gemma 4 31B IT is Google's largest instruction-tuned model in the Gemma line, designed for open, self-hosted chat and reasoning. It accepts image and text inputs and returns text, with a reasoning mode and tool calling. It has a 262,144-token context window and a 32,768-token output limit. Weights are open. Pricing starts at $0.09 per million input tokens and $0.25 per million output tokens, with cached input at $0.01 per million tokens. It is served by 35 hosts.
Specs
Pricing
First version of its line in the registry.
Range across the hosts that serve it: up to $0.99 per 1M input tokens.
Strengths
- 262,144-token context window
- 32,768-token max output
- Open weights for self-hosting
- Accepts image and text inputs
- Reasoning mode and tool calling
Best for
- Reach for it for self-hosted chat and reasoning
- Reach for it for processing long documents with a 262k context
- Reach for it for multimodal tasks combining image and text
- Reach for it for tool-calling workflows
How to access
35 hosts serve this model at the price above · the maker's documentation
Gemma: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 262,144 tokens.
- Are the weights open?
- Yes, the weights are open for self-hosting.
- What is the pricing?
- From $0.09 per million input tokens and $0.25 per million output tokens, with cached input at $0.01 per million tokens.