LLM Reference

Gemini 3.8 Flash

Released
2026-09-02
Last refreshed
2026-09-02
Status
Researched today
ProprietaryCommercial use: conditionalMultimodalCodingRAGAgentsLong contextVisionJSON / Tool use

Gemini 3.8 Flash is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, rag, and agents
  • Workloads that can use a 1.05m context window
  • Buyers comparing 2 tracked provider routes

Do not use it for

  • Workloads where another current model has stronger sourced task evidence
Specifications
Released
2026-09-02
Context
1.05m
Max output
65,536
Knowledge cutoff
2026-03
Specialization
general
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Created by

Pioneering artificial intelligence research.

London, United Kingdom
Founded 2014
Website
Pricing
Output / 1M
$3.75
Input / 1M
$0.750

Cheapest of 2 routes · Google AI Studio · cache read $0.075

About

Gemini 3.8 Flash is Google DeepMind's generally available most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows, released September 2, 2026. It accepts text, image, video, audio, and PDF inputs and returns text, with a 1,048,576-token context window, up to 65,536 output tokens, and Gemini API support for thinking (low/medium/high; minimal is not supported), function calling, tool use, structured outputs, code execution, prompt caching, search grounding, URL context, computer use (preview), and batch, flex, and priority consumption. Official model ID: gemini-3.8-flash. Sibling of live gemini-3.7-flash under a new family gemini-3.8.

Top use-case fit: coding, agents, and build tasks

Coding

Included by capability and metadata signals in the decision map.

RAG

Included by capability and metadata signals in the decision map.

Agents

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare all 2

Compare API pricing across 2 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MBatch in / outCacheRoute
Google AI Studio$0.750$3.75$0.375 / $1.88read $0.075
Serverless
OpenRouter$0.750$3.75$0.375 / $1.88read $0.075
Serverless

Available via routers & gateways(13)

Capabilities

VisionMultimodalReasoningJSON / Tool useStructured OutputsCode ExecutionPrompt CachingBatch APIAudio

Benchmark peer barsfor Coding

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

gemini-3.8-flash