25 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day

Model registry / Google / Gemini Flash / Gemini 3.8 Flash

Gemini 3.8 Flash

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows

Line

Gemini Flash

Weights

API only

Released

2026-09-02

Context

1.05M

Input

$0.75 / 1M

Output

$3.75 / 1M

Max output

66K

Coverage

4 signals

Overview

Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It is the latest in the gemini-flash line, succeeding Gemini 3.7 Flash. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. Pricing starts at $0.75 per million input tokens and $3.75 per million output tokens, with cached input at $0.075 per million tokens. The weights are not open. It supports reasoning mode and tool calling, and accepts audio, image, pdf, text, and video, returning text.

Specs

Released2026-09-02
LineGoogle · Gemini Flash
WeightsAPI only
Context1.05M tokens
Max output66K tokens
Inputaudio, image, pdf, text, video
Outputtext
ReasoningYes
Tool callingYes
Knowledge cutoffnot stated
API idgemini-3.8-flash

Pricing

Input$0.75 / 1M tokens
Cached input$0.075 / 1M tokens
Output$3.75 / 1M tokens

Same headline figures as Gemini 3.7 Flash.

Range across the hosts that serve it: up to $1.50 per 1M input tokens.

Strengths

  • Long-horizon software engineering and autonomous agents
  • 1M token context window for large workflows
  • Tool calling and reasoning mode included
  • Accepts audio, image, pdf, text, video
  • Priced from $0.75 per million input tokens

Best for

  • Reach for it for long-horizon software engineering tasks
  • Reach for it for building autonomous agents
  • Reach for it for complex enterprise workflows
  • Reach for it for processing multimodal inputs in one model

How to access

ProviderModel id
302.AIgemini-3.8-flash
AIHubMixgemini-3.8-flash
Cortecsgemini-3.8-flash
CrossModelgemini-3.8-flash
DevPass (LLM Gateway)gemini-3.8-flash
Eden AIgemini-3.8-flash
GitHub Copilotgemini-3.8-flash
Googlegemini-3.8-flash
Kilo Gatewaygemini-3.8-flash
LLM Gatewaygemini-3.8-flash
Merge Gatewaygemini-3.8-flash
NanoGPTgemini-3.8-flash

22 hosts serve this model at the price above · the maker's documentation

What the digest said

4 signals of the archive mention this model, each scored against the published bar. This list is rebuilt from the archive at every build, so it cannot go stale.

Gemini Flash: every version

VersionReleasedContextInput
Gemini 3.8 FlashCurrent2026-09-021.05M$0.75
Gemini 3.7 Flash2026-08-131.05M$0.75
Gemini 3.6 Flash2026-07-211.05M$0.375
Gemini 3.5 Flash2026-05-191.05M$0.1857
Gemini 3 Flash Preview2025-12-171.05M$0.07
Gemini 2.5 Flash2025-06-171.05M$0.09

The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.

FAQ

What is the context window and output limit?
Context window is 1,048,576 tokens, max output is 65,536 tokens.
Is the model open weights?
No, weights are not open.
What input types does it accept?
It accepts audio, image, pdf, text, and video, and returns text.
ProprietaryReasoning1.05M context