Model registry / IBM / Granite / Granite-4.0-H-Small
Granite-4.0-H-Small
Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads
Line
Granite
Weights
Open
Released
2025-10-02
Context
131K
Input
$0.0636 / 1M
Output
$0.265 / 1M
Max output
131K
Coverage
not covered yet
Overview
Granite-4.0-H-Small is an open-weight hybrid model from IBM for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads. It is the first version of its line in the registry, released on 2025-10-02. The model has a context window of 131072 tokens and a maximum output of 131072 tokens. It accepts text and returns text, supports tool calling, and does not have a reasoning mode. Weights are open. Pricing starts at $0.0636 per million input tokens and $0.265 per million output tokens. It is served by 1 host.
Specs
Pricing
First version of its line in the registry.
Strengths
- Open-weight hybrid model for enterprise tasks
- 131072-token context window
- 131072-token maximum output
- Supports tool calling
- Accepts and returns text only
Best for
- Reach for it for enterprise chat and coding tasks
- Reach for it for retrieval-augmented generation workloads
- Reach for it for tool-calling applications
- Reach for it for text-only generation with long context
How to access
1 host serve this model at the price above · the maker's documentation
Granite: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and output limit?
- The context window is 131072 tokens and the maximum output is also 131072 tokens.
- Does it support tool calling?
- Yes, it supports tool calling. It does not have a reasoning mode.
- What is the pricing?
- Pricing starts at $0.0636 per million input tokens and $0.265 per million output tokens.