Google Gemini model

Gemini 3.8 Flash

Google’s most intelligent Flash model for long-horizon software engineering, autonomous agents, complex enterprise workflows, and multimodal production systems that need frontier-level capability at Flash economics.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.75
intro · per 1M
Cached input
$0.075
intro · per 1M
Output
$3.75
intro · per 1M

01 / Overview

What Gemini 3.8 Flash Is

Gemini 3.8 Flash is Google’s current stable Flash model for workloads that need more than lightweight high-throughput inference: long-horizon coding, autonomous agents, critical multi-step reasoning, and complex enterprise execution.

Flash Has Become an Agentic Workhorse Tier

Google positions Gemini 3.8 Flash as its most intelligent Flash model.

That is an important distinction. The Flash name no longer means “use this only for simple or cheap tasks.” Gemini 3.8 Flash is explicitly engineered for long-running software engineering, autonomous agents, deterministic tool execution, and enterprise workflows where the model must maintain state across many decisions.

Google released gemini-3.8-flash as generally available on September 2, 2026.

The model accepts text, images, video, audio, and PDF input and generates text. It supports a 1,048,576-token input context window and up to 65,536 output tokens.

It also includes Google’s modern built-in tool stack: Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, and preview Computer Use.

  • Stable model ID: gemini-3.8-flash.
  • GA since September 2, 2026.
  • 1,048,576 input tokens.
  • 65,536 maximum output tokens.
  • Text, image, video, audio, and PDF input.
  • Text output.
  • Thinking levels: low, medium, high.
  • Default thinking level: medium.
Model profile
Provider
Google
Family
Gemini 3.8 Flash
Model ID
gemini-3.8-flash
Status
Stable · GA
Knowledge cutoff
Jan 2025
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / Software engineering

Built for Long-Horizon Software Engineering

Google identifies software engineering as one of the central improvements in Gemini 3.8 Flash, including real-world coding, complex multi-file refactoring, and deterministic tool execution.

Evaluate Whether the Model Can Finish the Whole Engineering Task

A coding benchmark can measure local correctness. A production coding agent must do considerably more.

It may need to inspect a repository, understand dependencies, plan a change, edit several files, call tools, run tests, diagnose failures, revise the implementation, and verify the final result.

That makes trajectory completion more important than isolated code generation quality.

For Gemini 3.8 Flash, useful engineering metrics include:

  • accepted implementation rate;
  • tests passed on first run;
  • number of tool calls;
  • repeated file reads or searches;
  • failed execution loops;
  • human interventions;
  • total thinking and output tokens;
  • latency to accepted result.

Google notes that 3.8 Flash can intentionally spend more tokens than 3.7 Flash on difficult long-running tasks because it takes smaller reasoning steps, calls tools iteratively, and verifies its work.

That can increase token consumption while still lowering the total cost of a successful task if it reduces failures and retries.

Long-horizon engineering
  1. 01

    Multi-file implementation

    Navigate repository structure, dependencies, source files, tests, and build configuration across a complete feature task.

  2. 02

    Complex refactoring

    Coordinate changes across multiple files while preserving behavior and validating the result with tools.

  3. 03

    Debugging loops

    Investigate failures, execute diagnostics, modify code, and re-test rather than stopping after the first hypothesis.

  4. 04

    Deterministic tool execution

    Evaluate whether tool calls, parameters, and execution order remain reliable across long coding trajectories.

03 / Autonomous agents

Autonomous Agents With More Reliable Multi-Step Execution

Gemini 3.8 Flash is explicitly designed for autonomous agents that need to plan, invoke tools, observe results, recover from errors, and continue toward a goal across many steps.

The Agent Is a Loop, Not a Single Completion

In agentic systems, model quality is partly an orchestration property.

The model must decide when to call a tool, which tool to use, what parameters to send, whether the result is sufficient, when to retry, and when to stop.

Google describes Gemini 3.8 Flash as reducing failed loops and errors compared with earlier Flash generations.

It is also the default model behind Google’s Antigravity managed agent, which reinforces its positioning as the primary agentic workhorse rather than a lightweight fallback model.

When evaluating agent behavior, measure:

  • successful completion without human rescue;
  • unnecessary loop count;
  • malformed function calls;
  • tool-selection accuracy;
  • recovery after failed tools;
  • total steps to completion;
  • total token consumption;
  • wall-clock latency.

A stronger model that uses more tokens can still be the cheaper agent if it reaches completion with fewer failed trajectories.

Agentic execution
  1. 01

    Plan

    Break a long objective into actionable steps while maintaining the original goal.

  2. 02

    Act

    Call built-in Google tools or application-defined functions with appropriate parameters.

  3. 03

    Observe

    Interpret tool results, multimodal evidence, errors, and intermediate state.

  4. 04

    Recover and verify

    Correct failed steps, verify the result, and stop when the task is actually complete.

04 / Pricing

Gemini 3.8 Flash Pricing

Gemini 3.8 Flash has introductory pricing through December 31, 2026, followed by higher standard rates beginning January 1, 2027.

Current Pricing Is Temporary

Through December 31, 2026, Google lists standard paid API pricing at:

  • $0.75 per 1M input tokens.
  • $3.75 per 1M output tokens, including thinking tokens.
  • $0.075 per 1M cached-context tokens.
  • $0.50 per 1M cached tokens per hour of storage.

Beginning January 1, 2027, the standard rates are scheduled to become:

  • $1.50 per 1M input tokens.
  • $7.50 per 1M output tokens.
  • $0.15 per 1M cached-context tokens.
  • $1.00 per 1M cached tokens per hour of storage.

Batch and Flex inference are priced at half the standard token rates. Priority inference is available at a premium.

This scheduled price transition matters for capacity planning. A product launched in late 2026 should not build its 2027 unit economics around the introductory rate.

Introductory and standard pricing

1M tokens · USD

Input
$0.75
Cached input
$0.075
Output + thinking
$3.75

Example: 20K input + 4K output

Input cost
$0.0150
Output cost
$0.0150
Estimated introductory total
$0.0300

05 / Multimodal context

A 1M-Token Multimodal Context Window

Gemini 3.8 Flash can process up to 1,048,576 input tokens across text, images, video, audio, and PDF content, with up to 65,536 output tokens.

The Context Window Is Not Limited to Text

This is one of the strongest differentiators of Gemini’s Flash tier.

A single workflow can combine a large codebase with screenshots, a PDF specification, recorded audio, video material, retrieved URLs, and tool results.

That makes context capacity useful for more than “very long prompts.” It becomes working memory for a multimodal agent.

Examples include:

  • reviewing a repository together with architecture PDFs;
  • analyzing user-interface screenshots while reading implementation code;
  • extracting evidence from long video or audio material;
  • comparing documents with live web-grounded information;
  • preserving long agent histories and tool outputs.

Large context still needs curation. Irrelevant media and historical state increase cost and can reduce the signal-to-noise ratio even when everything technically fits.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Input can include text, images, video, audio, and PDFs. The 1M-token window is shared across the complete request context.

06 / Thinking

Low, Medium, and High Thinking Levels

Gemini 3.8 Flash lets applications trade latency and token consumption for deeper reasoning through three supported thinking levels: low, medium, and high.

Medium Is the Default

Google sets medium as the default thinking level for Gemini 3.8 Flash.

The intended workload split is straightforward:

  • Low reduces time-to-answer for latency-sensitive work such as chat, incident-response pipelines, drafts, and fast analysis.
  • Medium is the default and recommended balance for most tasks, including complex code and agentic workflows.
  • High maximizes reasoning depth and tool orchestration for mathematics, deep analysis, and difficult multi-step tasks.

minimal is not supported on Gemini 3.8 Flash. Sending it returns an API validation error.

Thinking tokens are included in output-token billing, so raising the thinking level changes both latency and economics.

Google also notes that difficult 3.8 Flash tasks may consume more tokens than 3.7 Flash by design. The model may take more reasoning steps and perform more verification rather than optimizing for minimum token count.

Thinking levels
lowmedium · defaulthigh

Interactive workload

Use low thinking when responsiveness matters and evaluation shows that deeper reasoning does not materially improve the outcome.

Long-running agent

Use medium or high for difficult coding, tool orchestration, mathematics, and multi-step workflows where stronger verification improves completion.

07 / Built-in tools

A Broad Built-In Tool Surface

Gemini 3.8 Flash can operate as more than a multimodal completion model: Google exposes search, maps, files, code, URL retrieval, function calling, structured output, and computer interaction around the same model.

Tool Combination Matters for Production Agents

The official model specification lists support for:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • URL context;
  • function calling;
  • structured outputs;
  • context caching;
  • Computer Use in preview.

Google Search can provide current web information beyond the model’s January 2025 knowledge cutoff and return grounded citations.

Function calling lets applications expose their own operations. Built-in tools can also be combined with custom tools in Gemini 3 workflows.

Computer Use makes Gemini 3.8 Flash particularly relevant to GUI-based agents. Google currently recommends it as the model for Computer Use because of its UI interaction accuracy and tool-call reliability.

The practical evaluation question is not simply “does the model support tools?” It is whether the entire tool trajectory completes correctly with acceptable latency and cost.

Tool ecosystem
  • Google Search

    Ground responses in current web information and return source citations.

    Supported
  • Google Maps

    Use Maps grounding for location-aware workflows.

    Supported
  • Code execution

    Run generated code as part of reasoning and verification workflows.

    Supported
  • File Search

    Retrieve relevant content from indexed files.

    Supported
  • Function calling

    Invoke application-defined operations and combine them with built-in tools.

    Supported
  • Structured outputs

    Return machine-readable output that follows an expected schema.

    Supported
  • Computer Use

    Preview support for agents that interact with graphical user interfaces.

    Supported
  • URL context

    Read and reason over content from supplied URLs.

    Supported

08 / Agentic video

Agentic Video Understanding

Gemini 3.8 Flash can analyze long-form video using an agentic processing mode that dynamically decides which parts of the timeline, transcript, audio, and frames need closer inspection.

The Model Does Not Need to Process Every Frame Equally

Traditional static video processing samples frames at a fixed rate and places them into context.

Agentic video processing takes a different approach. The model can navigate the video timeline, retrieve only the relevant transcript segments or frames, and adjust inspection behavior based on the question.

Google reports that this mode can be substantially more token-efficient than static processing on long video while also improving quality.

This is useful for:

  • finding specific moments in long recordings;
  • analyzing lectures or meetings;
  • inspecting demonstrations and tutorials;
  • comparing several long videos;
  • extracting evidence from video archives.

For EidoStack-style evaluation, compare static and agentic processing on the same media task. Measure answer quality, input-token consumption, latency, and whether the model consistently retrieves the correct part of the video.

Agentic video understanding
  1. 01

    Navigate

    Inspect the timeline based on the task instead of processing every segment with the same fixed strategy.

  2. 02

    Retrieve

    Load transcript, frames, or audio selectively when they are relevant to the question.

  3. 03

    Reason

    Combine the selected evidence with text instructions and other multimodal context.

  4. 04

    Optimize

    Compare quality and token usage against static video processing for long-form workloads.

09 / 3.7 → 3.8

Migrating From Gemini 3.7 Flash to Gemini 3.8 Flash

Gemini 3.8 Flash preserves the 1M/64K Flash envelope and 2026 introductory price tier, but it changes the recommended reasoning and request configuration for applications moving from older Gemini integrations.

Do More Than Replace the Model ID

Google’s migration guidance recommends updating the target to gemini-3.8-flash and cleaning up several older request patterns.

For current integrations:

  • use thinking_level instead of thinking_budget;
  • do not set minimal for Gemini 3.8 Flash;
  • remove deprecated temperature, top_p, and top_k overrides from the 3.8 migration path;
  • remove candidate_count, which is unsupported in Gemini 3 and later;
  • remove prefilled model turns;
  • preserve correct multi-turn interaction state;
  • audit function-calling payloads and tool-result identifiers.

The performance tradeoff is also different.

Google states that Gemini 3.8 Flash delivers better accuracy and more reliable execution than 3.7 Flash, but can use more tokens—especially at higher effort levels.

That means a migration should compare completed-task economics, not just equal-token pricing.

Migration checklist
  1. 01

    Update the model ID

    Move production calls to gemini-3.8-flash and validate endpoint-specific request behavior.

  2. 02

    Migrate thinking controls

    Replace thinking_budget with low, medium, or high thinking_level and remove unsupported minimal.

  3. 03

    Clean request configuration

    Remove deprecated sampling overrides, candidate_count, prefilled model turns, and stale conversation assumptions.

  4. 04

    Re-baseline cost per task

    Measure whether higher token consumption is offset by fewer failed loops, better first-pass accuracy, and stronger completion reliability.

10 / Evaluation

Gemini 3.8 Flash Strengths and Limitations

Gemini 3.8 Flash is best evaluated as an agentic multimodal workhorse: stronger and more deliberate than earlier Flash models, but potentially more token-hungry on difficult tasks.

Strengths

  • Long-horizon engineering

    Designed for multi-file coding, iterative debugging, tool execution, and software tasks that continue across many steps.

  • Broad native multimodality

    Text, images, video, audio, and PDFs can share the same 1M-token working context.

  • Deep tool ecosystem

    Search, Maps, File Search, code execution, URL context, custom functions, structured outputs, Computer Use, and caching support complex agent architectures.

  • Flash price-performance

    Introductory $0.75/$3.75 pricing gives agentic workloads a substantially lower entry point than many premium frontier tiers.

What to consider

  • Introductory pricing expires

    Standard token rates are scheduled to double on January 1, 2027, so long-term unit economics need the post-intro price.

  • More tokens on difficult tasks

    Google explicitly notes that 3.8 Flash may use more tokens than 3.7 Flash as it reasons, verifies, and calls tools more extensively.

  • No minimal thinking

    The lowest supported thinking level is low; minimal returns a validation error.

  • Computer Use is preview

    GUI-agent support is available but remains a preview capability and should be isolated behind production safeguards and regression tests.

Evaluate the workhorse, not just the benchmark

Test Gemini 3.8 Flash on your real agent workflow

Replay coding, multimodal, tool-use, search, computer-use, and long-running agent tasks to compare accepted-result quality, thinking cost, tool-call efficiency, latency, and total tokens per completed task.

Start Free

Gemini 3.8 Flash can deliberately spend more tokens on difficult multi-step tasks, so production evaluation should measure cost per completed workflow rather than headline token price alone.

Common Questions

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is Google’s stable Flash model for long-horizon software engineering, autonomous agents, critical multi-step reasoning, and complex enterprise workflows. Google describes it as its most intelligent Flash model.

What is the Gemini 3.8 Flash model ID?

The stable Gemini API model ID is gemini-3.8-flash.

When was Gemini 3.8 Flash released?

Google announced Gemini 3.8 Flash as generally available on September 2, 2026.

What is the Gemini 3.8 Flash context window?

Gemini 3.8 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.8 Flash support?

The official model specification lists text, image, video, audio, and PDF as supported input types. The model produces text output.

What is the knowledge cutoff of Gemini 3.8 Flash?

Google documents January 2025 as the knowledge cutoff for Gemini 3 models. For current information, Gemini 3.8 Flash supports grounding with Google Search.

How much does Gemini 3.8 Flash cost?

Through December 31, 2026, standard paid pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Google plans to change standard pricing to $1.50 input and $7.50 output per million tokens on January 1, 2027.

How much does Gemini 3.8 Flash context caching cost?

During introductory pricing, cached context costs $0.075 per 1M tokens plus $0.50 per 1M tokens per hour of storage. Starting January 1, 2027, Google lists $0.15 per 1M cached tokens plus $1.00 per 1M tokens per hour of storage.

Does Gemini 3.8 Flash support thinking?

Yes. Supported thinking levels are low, medium, and high. Medium is the default.

Does Gemini 3.8 Flash support minimal thinking?

No. Google states that minimal is not supported by Gemini 3.8 Flash and returns an API validation error. Use low for the least reasoning among supported levels.

Does Gemini 3.8 Flash support images, audio, and video?

Yes. Gemini 3.8 Flash accepts image, audio, and video input in addition to text and PDFs. It also supports agentic video processing for long-form video analysis.

Does Gemini 3.8 Flash support Computer Use?

Yes, in preview. Google recommends Gemini 3.8 Flash for Computer Use and describes it as providing high-accuracy UI interaction and reliable tool calling.

What built-in tools does Gemini 3.8 Flash support?

Google lists Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use.

What is agentic video understanding?

Agentic video understanding lets Gemini dynamically navigate a video timeline and selectively inspect transcripts, frames, or audio instead of processing the entire video using a fixed sampling strategy.

How is Gemini 3.8 Flash different from Gemini 3.7 Flash?

Google positions 3.8 Flash as a substantial improvement in software engineering, autonomous agents, specialized multi-step reasoning, and execution reliability. It can also consume more tokens than 3.7 Flash on difficult tasks because it performs more reasoning and verification.

Should I use Gemini 3.8 Flash for a new application?

Gemini 3.8 Flash is the current stable Flash model to evaluate for new coding, multimodal, agentic, and enterprise workflows. Test it against your acceptance thresholds and model the January 2027 standard pricing if the application will run beyond the introductory period.

Model information

Last updated

Specifications, release status, pricing, thinking levels, multimodal input types, built-in tool support, computer use, agentic video behavior, and migration guidance on this page are based on official Google Gemini API and Google Cloud documentation.

Gemini 3.8 Flash — Pricing, 1M Context, Agentic Coding & Multimodal Tools | EidoStack