25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Google / Gemini Flash Lite / Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite

Fast Gemini model balancing multimodal reasoning, tool use, and cost

Line

Gemini Flash Lite

Weights

API only

Released

2026-07-21

Context

1.05M

Input

$0.15 / 1M

Output

$1.25 / 1M

Max output

66K

Coverage

not covered yet

Overview

Gemini 3.5 Flash Lite is a fast model from Google that balances multimodal reasoning, tool use, and cost. It is part of the gemini-flash-lite line, succeeding the Gemini 3.1 Flash Lite Preview. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. It accepts audio, image, pdf, text, and video, and returns text. It supports reasoning mode and tool calling. The weights are not open. Pricing starts at $0.15 per million input tokens and $1.25 per million output tokens, with cached input at $0.015 per million tokens.

Specs

Released2026-07-21
LineGoogle · Gemini Flash Lite
WeightsAPI only
Context1.05M tokens
Max output66K tokens
Inputaudio, image, pdf, text, video
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoff2026-03
API idgemini-3.5-flash-lite

Pricing

Input$0.15 / 1M tokens
Cached input$0.015 / 1M tokens
Output$1.25 / 1M tokens

Against Gemini 3.1 Flash Lite Preview: input $0.125 to $0.15, output $0.75 to $1.25, cached input $0.0125 to $0.015.

Range across the hosts that serve it: up to $0.33 per 1M input tokens.

Strengths

  • Handles audio, image, pdf, text, and video inputs
  • Supports reasoning mode and tool calling
  • Context window of 1,048,576 tokens
  • Output up to 65,536 tokens
  • Priced from $0.15 per million input tokens

Best for

  • Reach for it for multimodal reasoning across text, image, audio, and video
  • Reach for it for tool-calling workflows with large context
  • Reach for it for cost-sensitive applications needing high output

How to access

ProviderModel id
302.AIgemini-3.5-flash-lite
AIHubMixgemini-3.5-flash-lite
Abacusgemini-3.5-flash-lite
Cortecsgemini-3.5-flash-lite
CrossModelgemini-3.5-flash-lite
DevPass (LLM Gateway)gemini-3.5-flash-lite
Eden AIgemini-3.5-flash-lite
Googlegemini-3.5-flash-lite
Impossiblgemini-3.5-flash-lite
Kilo Gatewaygemini-3.5-flash-lite
LLM Gatewaygemini-3.5-flash-lite
Merge Gatewaygemini-3.5-flash-lite

24 hosts serve this model at the price above · the maker's documentation

Gemini Flash Lite: every version

VersionReleasedContextInput
Gemini 3.5 Flash LiteCurrent2026-07-211.05M$0.15
Gemini 3.1 Flash Lite Preview2026-03-031.05M$0.125
Gemini 2.5 Flash-Lite2025-06-171.05M$0.07

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window?
The context window is 1,048,576 tokens, and the maximum output is 65,536 tokens.
Does it support tool calling?
Yes, it supports tool calling and also has a reasoning mode. It accepts audio, image, pdf, text, and video, and returns text.
What is the price?
Input starts at $0.15 per million tokens, output at $1.25 per million tokens. Cached input is $0.015 per million tokens.
ProprietaryReasoning1.05M context