25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / IBM / Granite / Granite-4.0-H-Small

Granite-4.0-H-Small

Open-weight hybrid model for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads

Line

Granite

Weights

Open

Released

2025-10-02

Context

131K

Input

$0.0636 / 1M

Output

$0.265 / 1M

Max output

131K

Coverage

not covered yet

Overview

Granite-4.0-H-Small is an open-weight hybrid model from IBM for enterprise chat, coding, retrieval-augmented generation, and tool-calling workloads. It is the first version of its line in the registry, released on 2025-10-02. The model has a context window of 131072 tokens and a maximum output of 131072 tokens. It accepts text and returns text, supports tool calling, and does not have a reasoning mode. Weights are open. Pricing starts at $0.0636 per million input tokens and $0.265 per million output tokens. It is served by 1 host.

Specs

Released2025-10-02
LineIBM · Granite
WeightsOpen
Context131K tokens
Max output131K tokens
Inputtext
Outputtext
ReasoningNo
Tool callingYes
Knowledge cutoffnot stated
API idibm/granite-4-h-small

Pricing

Input$0.0636 / 1M tokens
Cached inputnot stated
Output$0.265 / 1M tokens

First version of its line in the registry.

Strengths

  • Open-weight hybrid model for enterprise tasks
  • 131072-token context window
  • 131072-token maximum output
  • Supports tool calling
  • Accepts and returns text only

Best for

  • Reach for it for enterprise chat and coding tasks
  • Reach for it for retrieval-augmented generation workloads
  • Reach for it for tool-calling applications
  • Reach for it for text-only generation with long context

How to access

ProviderModel id
watsonx.aiibm/granite-4-h-small

1 host serve this model at the price above · the maker's documentation

Granite: every version

VersionReleasedContextInput
Granite-4.0-H-SmallCurrent2025-10-02131K$0.0636

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and output limit?
The context window is 131072 tokens and the maximum output is also 131072 tokens.
Does it support tool calling?
Yes, it supports tool calling. It does not have a reasoning mode.
What is the pricing?
Pricing starts at $0.0636 per million input tokens and $0.265 per million output tokens.
Open weights131K context