Model registry / Google / Gemma / Gemma 4 26B A4B IT
Gemma 4 26B A4B IT
Open Gemma instruction model for efficient chat and self-hosted deployments
Line
Gemma
Weights
Open
Released
2026-04-02
Context
262K
Input
$0.042 / 1M
Output
$0.22 / 1M
Max output
33K
Coverage
not covered yet
Overview
Gemma 4 26B A4B IT is an open instruction model from Google for efficient chat and self-hosted deployments. It is the first version of the gemma line in this registry, released on 2026-04-02. The model has a context window of 262144 tokens and a maximum output of 32768 tokens. It accepts image and text inputs and returns text. It supports reasoning mode and tool calling. Weights are open. It is served by 21 hosts. Pricing starts at $0.042 per million input tokens and $0.22 per million output tokens, with cached input at $0.0075 per million tokens.
Specs
Pricing
First version of its line in the registry.
Range across the hosts that serve it: up to $0.25 per 1M input tokens.
Strengths
- 262k token context window
- 32k token max output
- Open weights for self-hosting
- Supports reasoning and tool calling
- Accepts image and text input
Best for
- Reach for it for efficient chat applications
- Reach for it for self-hosted deployments
- Reach for it for multimodal reasoning tasks
- Reach for it for tool-calling workflows
How to access
21 hosts serve this model at the price above · the maker's documentation
Gemma: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window size?
- The context window is 262144 tokens, and the maximum output is 32768 tokens.
- Is the model open weights?
- Yes, the weights are open, and the model is served by 21 hosts.
- What does it cost to use?
- Pricing starts at $0.042 per million input tokens and $0.22 per million output tokens. Cached input is $0.0075 per million tokens.