Model registry / Google / Gemini Flash / Gemini 3.5 Flash
Gemini 3.5 Flash
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Line
Gemini Flash
Weights
API only
Released
2026-05-19
Context
1.05M
Input
$0.1857 / 1M
Output
$1.11 / 1M
Max output
66K
Coverage
1 signal
Overview
Gemini 3.5 Flash is a fast model in the gemini-flash line, balancing multimodal reasoning, tool use, and cost. It is the successor to the Gemini 3 Flash Preview, released on 2026-05-19. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. Pricing starts at $0.1857 per million input tokens and $1.11 per million output tokens, with cached input at $0.075 per million tokens. The weights are not open. It accepts audio, image, pdf, text, and video, and returns text. It supports reasoning mode and tool calling. Knowledge cutoff is January 2025.
Specs
Pricing
Against Gemini 3 Flash Preview: input $0.07 to $0.1857, output $0.43 to $1.11, cached input $0.025 to $0.075.
Range across the hosts that serve it: up to $1.65 per 1M input tokens.
Strengths
- 1,048,576 token context window
- 65,536 token max output
- Accepts audio, image, pdf, text, video
- Supports reasoning and tool calling
- Returns text only
Best for
- Reach for it for multimodal reasoning with large context
- Reach for it for tool-calling workflows
- Reach for it for processing audio, image, and video inputs
- Reach for it for high-volume text generation with 64K output
How to access
31 hosts serve this model at the price above · the maker's documentation
What the digest said
1 signal of the archive mention this model, each scored against the published bar. This list is rebuilt from the archive at every build, so it cannot go stale.
Gemini Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 1,048,576 tokens.
- Does it support tool calling?
- Yes, it supports tool calling and reasoning mode.
- What is the pricing?
- Input from $0.1857 per million tokens, output from $1.11 per million tokens, cached input from $0.075 per million tokens.