Model registry / Google / Gemini Flash Lite / Gemini 3.1 Flash Lite Preview
Gemini 3.1 Flash Lite Preview
Legacy model retained for compatibility with older integrations
Line
Gemini Flash Lite
Weights
API only
Released
2026-03-03
Context
1.05M
Input
$0.125 / 1M
Output
$0.75 / 1M
Max output
66K
Coverage
not covered yet
Overview
Gemini 3.1 Flash Lite Preview is a legacy model in the gemini-flash-lite line, retained by Google for compatibility with older integrations. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. Pricing starts at $0.125 per million input tokens and $0.75 per million output tokens, with cached input at $0.0125 per million tokens. The weights are not open. It supports reasoning mode and tool calling, and accepts audio, image, pdf, text, and video inputs, returning text. Knowledge cutoff is January 2025.
Specs
Pricing
Against Gemini 2.5 Flash-Lite: input $0.07 to $0.125, output $0.10 to $0.75, cached input $0.01 to $0.0125.
Range across the hosts that serve it: up to $0.272 per 1M input tokens.
Strengths
- 1,048,576 token context window
- 65,536 token maximum output
- Supports reasoning mode
- Supports tool calling
- Accepts audio, image, pdf, text, video
Best for
- Reach for it for maintaining compatibility with older integrations
- Reach for it for processing very long documents up to 1M tokens
- Reach for it for multimodal analysis with audio, image, pdf, text, video
How to access
25 hosts serve this model at the price above · the maker's documentation
Gemini Flash Lite: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window?
- The context window is 1,048,576 tokens.
- What is the pricing?
- Input from $0.125 per million tokens, output from $0.75 per million tokens, cached input from $0.0125 per million tokens.
- What input types does it accept?
- It accepts audio, image, pdf, text, and video, and returns text.