Model registry / Google / Gemini Flash Lite / Gemini 2.5 Flash-Lite
Gemini 2.5 Flash-Lite
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Line
Gemini Flash Lite
Weights
API only
Released
2025-06-17
Context
1.05M
Input
$0.07 / 1M
Output
$0.10 / 1M
Max output
66K
Coverage
not covered yet
Overview
Gemini 2.5 Flash-Lite is a lean version of Gemini 2.5 for cheap multimodal traffic and quick agents. It is the first version of its line in the registry, released on 2025-06-17. The context window is 1,048,576 tokens and the maximum output is 65,536 tokens. Price starts at $0.07 per million input tokens and $0.10 per million output tokens, with cached input at $0.01 per million tokens. The weights are not open. It accepts audio, image, pdf, text, and video, and returns text. It supports reasoning mode and tool calling.
Specs
Pricing
First version of its line in the registry.
Range across the hosts that serve it: up to $0.10 per 1M input tokens.
Strengths
- 1M token context window
- 65K token max output
- Multimodal input: audio, image, pdf, text, video
- Reasoning mode and tool calling
- Low price from $0.07 per M input
Best for
- Reach for it for cheap multimodal traffic
- Reach for it for quick agents
- Reach for it for large context tasks
- Reach for it for tool calling workflows
How to access
24 hosts serve this model at the price above · the maker's documentation
Gemini Flash Lite: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window size?
- The context window is 1,048,576 tokens.
- Does it support image input?
- Yes, it accepts audio, image, pdf, text, and video.
- Are the weights open?
- No, the weights are not open.