Model registry / Google / Gemini Flash / Gemini 3.8 Flash
Gemini 3.8 Flash
Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows
Line
Gemini Flash
Weights
API only
Released
2026-09-02
Context
1.05M
Input
$0.75 / 1M
Output
$3.75 / 1M
Max output
66K
Coverage
4 signals
Overview
Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It is the latest in the gemini-flash line, succeeding Gemini 3.7 Flash. It has a context window of 1,048,576 tokens and a maximum output of 65,536 tokens. Pricing starts at $0.75 per million input tokens and $3.75 per million output tokens, with cached input at $0.075 per million tokens. The weights are not open. It supports reasoning mode and tool calling, and accepts audio, image, pdf, text, and video, returning text.
Specs
Pricing
Same headline figures as Gemini 3.7 Flash.
Range across the hosts that serve it: up to $1.50 per 1M input tokens.
Strengths
- Long-horizon software engineering and autonomous agents
- 1M token context window for large workflows
- Tool calling and reasoning mode included
- Accepts audio, image, pdf, text, video
- Priced from $0.75 per million input tokens
Best for
- Reach for it for long-horizon software engineering tasks
- Reach for it for building autonomous agents
- Reach for it for complex enterprise workflows
- Reach for it for processing multimodal inputs in one model
How to access
22 hosts serve this model at the price above · the maker's documentation
What the digest said
- 8[AINews] Opus 5.5 is good at explainer videos2026-09-29
- 8[AINews] The Future of Latent Space2026-09-26
- 9Gemini 3.8 TTS Playground2026-09-24
- 8Qwen 发布原生全模态模型 Qwen3.8-Omni-Flash,主打音视频智能体任务交付2026-09-18
4 signals of the archive mention this model, each scored against the published bar. This list is rebuilt from the archive at every build, so it cannot go stale.
Gemini Flash: every version
The registry keeps the six most recent versions of a line. Older ones stay in the file and can be surfaced without a rebuild.
FAQ
- What is the context window and output limit?
- Context window is 1,048,576 tokens, max output is 65,536 tokens.
- Is the model open weights?
- No, weights are not open.
- What input types does it accept?
- It accepts audio, image, pdf, text, and video, and returns text.