Google Gemini model
Gemini 3.8 Flash
Google’s most intelligent Flash model for long-horizon software engineering, autonomous agents, complex enterprise workflows, and multimodal production systems that need frontier-level capability at Flash economics.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $0.75
- intro · per 1M
- Cached input
- $0.075
- intro · per 1M
- Output
- $3.75
- intro · per 1M
01 / Overview
What Gemini 3.8 Flash Is
Gemini 3.8 Flash is Google’s current stable Flash model for workloads that need more than lightweight high-throughput inference: long-horizon coding, autonomous agents, critical multi-step reasoning, and complex enterprise execution.
Flash Has Become an Agentic Workhorse Tier
Google positions Gemini 3.8 Flash as its most intelligent Flash model.
That is an important distinction. The Flash name no longer means “use this only for simple or cheap tasks.” Gemini 3.8 Flash is explicitly engineered for long-running software engineering, autonomous agents, deterministic tool execution, and enterprise workflows where the model must maintain state across many decisions.
Google released gemini-3.8-flash as generally available on September 2, 2026.
The model accepts text, images, video, audio, and PDF input and generates text. It supports a 1,048,576-token input context window and up to 65,536 output tokens.
It also includes Google’s modern built-in tool stack: Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, and preview Computer Use.
- Stable model ID:
gemini-3.8-flash. - GA since September 2, 2026.
- 1,048,576 input tokens.
- 65,536 maximum output tokens.
- Text, image, video, audio, and PDF input.
- Text output.
- Thinking levels: low, medium, high.
- Default thinking level: medium.
- Provider
- Family
- Gemini 3.8 Flash
- Model ID
- gemini-3.8-flash
- Status
- Stable · GA
- Knowledge cutoff
- Jan 2025
- Input
- Text, Image, Video, Audio, PDF
- Output
- Text
02 / Software engineering
Built for Long-Horizon Software Engineering
Google identifies software engineering as one of the central improvements in Gemini 3.8 Flash, including real-world coding, complex multi-file refactoring, and deterministic tool execution.
Evaluate Whether the Model Can Finish the Whole Engineering Task
A coding benchmark can measure local correctness. A production coding agent must do considerably more.
It may need to inspect a repository, understand dependencies, plan a change, edit several files, call tools, run tests, diagnose failures, revise the implementation, and verify the final result.
That makes trajectory completion more important than isolated code generation quality.
For Gemini 3.8 Flash, useful engineering metrics include:
- accepted implementation rate;
- tests passed on first run;
- number of tool calls;
- repeated file reads or searches;
- failed execution loops;
- human interventions;
- total thinking and output tokens;
- latency to accepted result.
Google notes that 3.8 Flash can intentionally spend more tokens than 3.7 Flash on difficult long-running tasks because it takes smaller reasoning steps, calls tools iteratively, and verifies its work.
That can increase token consumption while still lowering the total cost of a successful task if it reduces failures and retries.
- 01
Multi-file implementation
Navigate repository structure, dependencies, source files, tests, and build configuration across a complete feature task.
- 02
Complex refactoring
Coordinate changes across multiple files while preserving behavior and validating the result with tools.
- 03
Debugging loops
Investigate failures, execute diagnostics, modify code, and re-test rather than stopping after the first hypothesis.
- 04
Deterministic tool execution
Evaluate whether tool calls, parameters, and execution order remain reliable across long coding trajectories.
03 / Autonomous agents
Autonomous Agents With More Reliable Multi-Step Execution
Gemini 3.8 Flash is explicitly designed for autonomous agents that need to plan, invoke tools, observe results, recover from errors, and continue toward a goal across many steps.
The Agent Is a Loop, Not a Single Completion
In agentic systems, model quality is partly an orchestration property.
The model must decide when to call a tool, which tool to use, what parameters to send, whether the result is sufficient, when to retry, and when to stop.
Google describes Gemini 3.8 Flash as reducing failed loops and errors compared with earlier Flash generations.
It is also the default model behind Google’s Antigravity managed agent, which reinforces its positioning as the primary agentic workhorse rather than a lightweight fallback model.
When evaluating agent behavior, measure:
- successful completion without human rescue;
- unnecessary loop count;
- malformed function calls;
- tool-selection accuracy;
- recovery after failed tools;
- total steps to completion;
- total token consumption;
- wall-clock latency.
A stronger model that uses more tokens can still be the cheaper agent if it reaches completion with fewer failed trajectories.
- 01
Plan
Break a long objective into actionable steps while maintaining the original goal.
- 02
Act
Call built-in Google tools or application-defined functions with appropriate parameters.
- 03
Observe
Interpret tool results, multimodal evidence, errors, and intermediate state.
- 04
Recover and verify
Correct failed steps, verify the result, and stop when the task is actually complete.
04 / Pricing
Gemini 3.8 Flash Pricing
Gemini 3.8 Flash has introductory pricing through December 31, 2026, followed by higher standard rates beginning January 1, 2027.
Current Pricing Is Temporary
Through December 31, 2026, Google lists standard paid API pricing at:
- $0.75 per 1M input tokens.
- $3.75 per 1M output tokens, including thinking tokens.
- $0.075 per 1M cached-context tokens.
- $0.50 per 1M cached tokens per hour of storage.
Beginning January 1, 2027, the standard rates are scheduled to become:
- $1.50 per 1M input tokens.
- $7.50 per 1M output tokens.
- $0.15 per 1M cached-context tokens.
- $1.00 per 1M cached tokens per hour of storage.
Batch and Flex inference are priced at half the standard token rates. Priority inference is available at a premium.
This scheduled price transition matters for capacity planning. A product launched in late 2026 should not build its 2027 unit economics around the introductory rate.
1M tokens · USD
- Input
- $0.75
- Cached input
- $0.075
- Output + thinking
- $3.75
Example: 20K input + 4K output
- Input cost
- $0.0150
- Output cost
- $0.0150
- Estimated introductory total
- $0.0300
05 / Multimodal context
A 1M-Token Multimodal Context Window
Gemini 3.8 Flash can process up to 1,048,576 input tokens across text, images, video, audio, and PDF content, with up to 65,536 output tokens.
The Context Window Is Not Limited to Text
This is one of the strongest differentiators of Gemini’s Flash tier.
A single workflow can combine a large codebase with screenshots, a PDF specification, recorded audio, video material, retrieved URLs, and tool results.
That makes context capacity useful for more than “very long prompts.” It becomes working memory for a multimodal agent.
Examples include:
- reviewing a repository together with architecture PDFs;
- analyzing user-interface screenshots while reading implementation code;
- extracting evidence from long video or audio material;
- comparing documents with live web-grounded information;
- preserving long agent histories and tool outputs.
Large context still needs curation. Irrelevant media and historical state increase cost and can reduce the signal-to-noise ratio even when everything technically fits.
Input context
1,048,576
Max output
65,536
Input can include text, images, video, audio, and PDFs. The 1M-token window is shared across the complete request context.
06 / Thinking
Low, Medium, and High Thinking Levels
Gemini 3.8 Flash lets applications trade latency and token consumption for deeper reasoning through three supported thinking levels: low, medium, and high.
Medium Is the Default
Google sets medium as the default thinking level for Gemini 3.8 Flash.
The intended workload split is straightforward:
- Low reduces time-to-answer for latency-sensitive work such as chat, incident-response pipelines, drafts, and fast analysis.
- Medium is the default and recommended balance for most tasks, including complex code and agentic workflows.
- High maximizes reasoning depth and tool orchestration for mathematics, deep analysis, and difficult multi-step tasks.
minimal is not supported on Gemini 3.8 Flash. Sending it returns an API validation error.
Thinking tokens are included in output-token billing, so raising the thinking level changes both latency and economics.
Google also notes that difficult 3.8 Flash tasks may consume more tokens than 3.7 Flash by design. The model may take more reasoning steps and perform more verification rather than optimizing for minimum token count.
Interactive workload
Use low thinking when responsiveness matters and evaluation shows that deeper reasoning does not materially improve the outcome.
Long-running agent
Use medium or high for difficult coding, tool orchestration, mathematics, and multi-step workflows where stronger verification improves completion.
07 / Built-in tools
A Broad Built-In Tool Surface
Gemini 3.8 Flash can operate as more than a multimodal completion model: Google exposes search, maps, files, code, URL retrieval, function calling, structured output, and computer interaction around the same model.
Tool Combination Matters for Production Agents
The official model specification lists support for:
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- URL context;
- function calling;
- structured outputs;
- context caching;
- Computer Use in preview.
Google Search can provide current web information beyond the model’s January 2025 knowledge cutoff and return grounded citations.
Function calling lets applications expose their own operations. Built-in tools can also be combined with custom tools in Gemini 3 workflows.
Computer Use makes Gemini 3.8 Flash particularly relevant to GUI-based agents. Google currently recommends it as the model for Computer Use because of its UI interaction accuracy and tool-call reliability.
The practical evaluation question is not simply “does the model support tools?” It is whether the entire tool trajectory completes correctly with acceptable latency and cost.
- Supported
Google Search
Ground responses in current web information and return source citations.
- Supported
Google Maps
Use Maps grounding for location-aware workflows.
- Supported
Code execution
Run generated code as part of reasoning and verification workflows.
- Supported
File Search
Retrieve relevant content from indexed files.
- Supported
Function calling
Invoke application-defined operations and combine them with built-in tools.
- Supported
Structured outputs
Return machine-readable output that follows an expected schema.
- Supported
Computer Use
Preview support for agents that interact with graphical user interfaces.
- Supported
URL context
Read and reason over content from supplied URLs.
08 / Agentic video
Agentic Video Understanding
Gemini 3.8 Flash can analyze long-form video using an agentic processing mode that dynamically decides which parts of the timeline, transcript, audio, and frames need closer inspection.
The Model Does Not Need to Process Every Frame Equally
Traditional static video processing samples frames at a fixed rate and places them into context.
Agentic video processing takes a different approach. The model can navigate the video timeline, retrieve only the relevant transcript segments or frames, and adjust inspection behavior based on the question.
Google reports that this mode can be substantially more token-efficient than static processing on long video while also improving quality.
This is useful for:
- finding specific moments in long recordings;
- analyzing lectures or meetings;
- inspecting demonstrations and tutorials;
- comparing several long videos;
- extracting evidence from video archives.
For EidoStack-style evaluation, compare static and agentic processing on the same media task. Measure answer quality, input-token consumption, latency, and whether the model consistently retrieves the correct part of the video.
- 01
Navigate
Inspect the timeline based on the task instead of processing every segment with the same fixed strategy.
- 02
Retrieve
Load transcript, frames, or audio selectively when they are relevant to the question.
- 03
Reason
Combine the selected evidence with text instructions and other multimodal context.
- 04
Optimize
Compare quality and token usage against static video processing for long-form workloads.
09 / 3.7 → 3.8
Migrating From Gemini 3.7 Flash to Gemini 3.8 Flash
Gemini 3.8 Flash preserves the 1M/64K Flash envelope and 2026 introductory price tier, but it changes the recommended reasoning and request configuration for applications moving from older Gemini integrations.
Do More Than Replace the Model ID
Google’s migration guidance recommends updating the target to gemini-3.8-flash and cleaning up several older request patterns.
For current integrations:
- use
thinking_levelinstead ofthinking_budget; - do not set
minimalfor Gemini 3.8 Flash; - remove deprecated
temperature,top_p, andtop_koverrides from the 3.8 migration path; - remove
candidate_count, which is unsupported in Gemini 3 and later; - remove prefilled model turns;
- preserve correct multi-turn interaction state;
- audit function-calling payloads and tool-result identifiers.
The performance tradeoff is also different.
Google states that Gemini 3.8 Flash delivers better accuracy and more reliable execution than 3.7 Flash, but can use more tokens—especially at higher effort levels.
That means a migration should compare completed-task economics, not just equal-token pricing.
- 01
Update the model ID
Move production calls to gemini-3.8-flash and validate endpoint-specific request behavior.
- 02
Migrate thinking controls
Replace thinking_budget with low, medium, or high thinking_level and remove unsupported minimal.
- 03
Clean request configuration
Remove deprecated sampling overrides, candidate_count, prefilled model turns, and stale conversation assumptions.
- 04
Re-baseline cost per task
Measure whether higher token consumption is offset by fewer failed loops, better first-pass accuracy, and stronger completion reliability.
10 / Evaluation
Gemini 3.8 Flash Strengths and Limitations
Gemini 3.8 Flash is best evaluated as an agentic multimodal workhorse: stronger and more deliberate than earlier Flash models, but potentially more token-hungry on difficult tasks.
Strengths
Long-horizon engineering
Designed for multi-file coding, iterative debugging, tool execution, and software tasks that continue across many steps.
Broad native multimodality
Text, images, video, audio, and PDFs can share the same 1M-token working context.
Deep tool ecosystem
Search, Maps, File Search, code execution, URL context, custom functions, structured outputs, Computer Use, and caching support complex agent architectures.
Flash price-performance
Introductory $0.75/$3.75 pricing gives agentic workloads a substantially lower entry point than many premium frontier tiers.
What to consider
Introductory pricing expires
Standard token rates are scheduled to double on January 1, 2027, so long-term unit economics need the post-intro price.
More tokens on difficult tasks
Google explicitly notes that 3.8 Flash may use more tokens than 3.7 Flash as it reasons, verifies, and calls tools more extensively.
No minimal thinking
The lowest supported thinking level is low; minimal returns a validation error.
Computer Use is preview
GUI-agent support is available but remains a preview capability and should be isolated behind production safeguards and regression tests.
Evaluate the workhorse, not just the benchmark
Test Gemini 3.8 Flash on your real agent workflow
Replay coding, multimodal, tool-use, search, computer-use, and long-running agent tasks to compare accepted-result quality, thinking cost, tool-call efficiency, latency, and total tokens per completed task.
Start FreeGemini 3.8 Flash can deliberately spend more tokens on difficult multi-step tasks, so production evaluation should measure cost per completed workflow rather than headline token price alone.
Common Questions
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google’s stable Flash model for long-horizon software engineering, autonomous agents, critical multi-step reasoning, and complex enterprise workflows. Google describes it as its most intelligent Flash model.
What is the Gemini 3.8 Flash model ID?
The stable Gemini API model ID is gemini-3.8-flash.
When was Gemini 3.8 Flash released?
Google announced Gemini 3.8 Flash as generally available on September 2, 2026.
What is the Gemini 3.8 Flash context window?
Gemini 3.8 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3.8 Flash support?
The official model specification lists text, image, video, audio, and PDF as supported input types. The model produces text output.
What is the knowledge cutoff of Gemini 3.8 Flash?
Google documents January 2025 as the knowledge cutoff for Gemini 3 models. For current information, Gemini 3.8 Flash supports grounding with Google Search.
How much does Gemini 3.8 Flash cost?
Through December 31, 2026, standard paid pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Google plans to change standard pricing to $1.50 input and $7.50 output per million tokens on January 1, 2027.
How much does Gemini 3.8 Flash context caching cost?
During introductory pricing, cached context costs $0.075 per 1M tokens plus $0.50 per 1M tokens per hour of storage. Starting January 1, 2027, Google lists $0.15 per 1M cached tokens plus $1.00 per 1M tokens per hour of storage.
Does Gemini 3.8 Flash support thinking?
Yes. Supported thinking levels are low, medium, and high. Medium is the default.
Does Gemini 3.8 Flash support minimal thinking?
No. Google states that minimal is not supported by Gemini 3.8 Flash and returns an API validation error. Use low for the least reasoning among supported levels.
Does Gemini 3.8 Flash support images, audio, and video?
Yes. Gemini 3.8 Flash accepts image, audio, and video input in addition to text and PDFs. It also supports agentic video processing for long-form video analysis.
Does Gemini 3.8 Flash support Computer Use?
Yes, in preview. Google recommends Gemini 3.8 Flash for Computer Use and describes it as providing high-accuracy UI interaction and reliable tool calling.
What built-in tools does Gemini 3.8 Flash support?
Google lists Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use.
What is agentic video understanding?
Agentic video understanding lets Gemini dynamically navigate a video timeline and selectively inspect transcripts, frames, or audio instead of processing the entire video using a fixed sampling strategy.
How is Gemini 3.8 Flash different from Gemini 3.7 Flash?
Google positions 3.8 Flash as a substantial improvement in software engineering, autonomous agents, specialized multi-step reasoning, and execution reliability. It can also consume more tokens than 3.7 Flash on difficult tasks because it performs more reasoning and verification.
Should I use Gemini 3.8 Flash for a new application?
Gemini 3.8 Flash is the current stable Flash model to evaluate for new coding, multimodal, agentic, and enterprise workflows. Test it against your acceptance thresholds and model the January 2027 standard pricing if the application will run beyond the introductory period.
Model information
Last updated
Specifications, release status, pricing, thinking levels, multimodal input types, built-in tool support, computer use, agentic video behavior, and migration guidance on this page are based on official Google Gemini API and Google Cloud documentation.