Model registry / Google / Gemini Flash Lite / Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Line
Gemini Flash Lite
Weights
API only
Released
2026-07-21
Context
1.05M
Input
$0.15 / 1M
Output
$1.25 / 1M
Max output
66K
Coverage
not covered yet
Overview
Gemini 3.5 Flash Lite is a fast model from Google that balances multimodal reasoning, tool use, and cost. It is part of the gemini-flash-lite line, succeeding the Gemini 3.1 Flash Lite Preview. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. It accepts audio, image, pdf, text, and video, and returns text. It supports reasoning mode and tool calling. The weights are not open. Pricing starts at $0.15 per million input tokens and $1.25 per million output tokens, with cached input at $0.015 per million tokens.
Specs
Pricing
Against Gemini 3.1 Flash Lite Preview: input $0.125 to $0.15, output $0.75 to $1.25, cached input $0.0125 to $0.015.
Range across the hosts that serve it: up to $0.33 per 1M input tokens.
Strengths
- Handles audio, image, pdf, text, and video inputs
- Supports reasoning mode and tool calling
- Context window of 1,048,576 tokens
- Output up to 65,536 tokens
- Priced from $0.15 per million input tokens
Best for
- Reach for it for multimodal reasoning across text, image, audio, and video
- Reach for it for tool-calling workflows with large context
- Reach for it for cost-sensitive applications needing high output
How to access
24 hosts serve this model at the price above · the maker's documentation
Gemini Flash Lite: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 1,048,576 tokens, and the maximum output is 65,536 tokens.
- Does it support tool calling?
- Yes, it supports tool calling and also has a reasoning mode. It accepts audio, image, pdf, text, and video, and returns text.
- What is the price?
- Input starts at $0.15 per million tokens, output at $1.25 per million tokens. Cached input is $0.015 per million tokens.