Model registry / Google / Gemini Flash / Gemini 2.5 Flash
Gemini 2.5 Flash
Fast Gemini workhorse for multimodal apps where latency and price matter
Line
Gemini Flash
Weights
API only
Released
2025-06-17
Context
1.05M
Input
$0.09 / 1M
Output
$0.71 / 1M
Max output
66K
Coverage
1 signal
Overview
Gemini 2.5 Flash is a fast multimodal model for applications where latency and price matter. It is the first version of the gemini-flash line in the registry. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. It accepts audio, image, pdf, text, and video, and returns text. It supports reasoning mode and tool calling. Weights are not open. Pricing starts at $0.09 per million input tokens and $0.71 per million output tokens, with cached input at $0.021 per million tokens.
Specs
Pricing
First version of its line in the registry.
Range across the hosts that serve it: up to $0.30 per 1M input tokens.
Strengths
- 1,048,576 token context window
- 65,536 token max output
- Accepts audio, image, pdf, text, video
- Supports reasoning mode and tool calling
- Priced from $0.09 per million input tokens
Best for
- Reach for it for multimodal apps needing low latency and low price
- Reach for it for processing long documents or videos within a 1M token window
- Reach for it for building agents with tool calling and reasoning
- Reach for it for generating text from audio, image, pdf, or video inputs
How to access
31 hosts serve this model at the price above · the maker's documentation
What the digest said
1 signal of the archive mention this model, each scored against the published bar. This list is rebuilt from the archive at every build, so it cannot go stale.
Gemini Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- 1,048,576 tokens.
- What are the input and output prices?
- From $0.09 per million input tokens and $0.71 per million output tokens. Cached input is $0.021 per million tokens.
- Does it support tool calling?
- Yes, it supports tool calling and reasoning mode.