Google Gemini stable model

Gemini 3.5 Flash

The first Gemini 3.5 Flash model to combine frontier-level intelligence with fast agentic action for coding, long-horizon workflows, multimodal reasoning, and computer-use automation.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$1.50
per 1M tokens
Cached input
$0.15
per 1M tokens
Output
$9.00
per 1M tokens

01 / Overview

What Gemini 3.5 Flash Is

Gemini 3.5 Flash is Google’s stable Flash model that brought frontier-level coding, agentic execution, and long-horizon reasoning into the high-speed Gemini Flash tier.

The First Gemini 3.5 Model

Google introduced Gemini 3.5 Flash on May 19, 2026 as the first model in the Gemini 3.5 family.

The launch theme was frontier intelligence with action.

That positioning matters because Gemini 3.5 Flash was not introduced as a lightweight fallback for simple requests. Google explicitly targeted difficult coding, autonomous agents, multi-step workflows, sub-agent deployment, and long-horizon tasks.

Google also reported that the model outperformed Gemini 3.1 Pro on several coding and agentic benchmarks while running substantially faster than other frontier models.

The stable Gemini API model ID is gemini-3.5-flash.

It accepts text, images, video, audio, and PDFs and generates text. The model supports up to 1,048,576 input tokens and up to 65,536 output tokens.

Gemini 3.5 Flash also supports dynamic thinking, automatic thought preservation, Google’s built-in tool ecosystem, context caching, and Computer Use in preview.

  • Stable model ID: gemini-3.5-flash.
  • Released May 19, 2026.
  • Stable and GA for scaled production use.
  • 1,048,576 input tokens.
  • 65,536 maximum output tokens.
  • Text, image, video, audio, and PDF input.
  • Text output.
  • Default thinking level: medium.
Model profile
Provider
Google
Family
Gemini 3.5 Flash
Model ID
gemini-3.5-flash
Status
Stable · GA
Knowledge cutoff
Jan 2025
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / Frontier + action

Frontier Intelligence at Flash Speed

Gemini 3.5 Flash was designed around a specific product idea: combine high-end reasoning capability with the responsiveness and throughput traditionally associated with Flash.

Quality and Latency Are Both Part of the Model

At launch, Google described 3.5 Flash as delivering intelligence comparable with large flagship models while preserving Flash-class speed.

Google also reported roughly four-times-higher output-token throughput than other frontier models in its launch comparisons.

That makes the model relevant to applications where a strong result is not enough if it arrives too slowly.

Examples include:

  • interactive coding agents;
  • browser automation;
  • real-time operational analysis;
  • multi-step customer workflows;
  • document-heavy productivity tools;
  • agent orchestration where one model call triggers several downstream actions.

For these systems, evaluate both task quality and time to useful completion.

A model can have excellent benchmark scores but still be the wrong product choice if long reasoning delays every user-visible action.

Gemini 3.5 Flash’s SEO and evaluation identity therefore sits between two older assumptions: it is more capable than a traditional speed-first model, but less expensive and faster than a Pro-tier frontier model.

Frontier intelligence with action
  1. 01

    Reason

    Handle difficult coding, multimodal, and multi-step problems with frontier-level capability.

  2. 02

    Act

    Move from reasoning into tools, sub-agents, browser interaction, code execution, and application-defined functions.

  3. 03

    Iterate quickly

    Use Flash-class responsiveness for workflows that require repeated reason-act-observe loops.

  4. 04

    Measure completion latency

    Track total time to a successful outcome instead of comparing isolated response latency.

03 / Agentic coding

Built for Rapid Agentic Coding Loops

Google describes Gemini 3.5 Flash as particularly effective for rapid agentic loops involving complex coding cycles, iterative exploration, and long-running software tasks.

Coding Quality Is a Trajectory Property

Repository-level engineering rarely succeeds in one generation.

A coding agent may need to:

  1. inspect source files;
  2. understand dependencies;
  3. form a plan;
  4. modify several files;
  5. execute tests;
  6. diagnose failures;
  7. search for more context;
  8. repair the implementation;
  9. verify the final state.

Gemini 3.5 Flash was explicitly built for this loop.

Google’s launch materials emphasized coding, sub-agent deployment, long-horizon task execution, and rapid iteration over alternate approaches.

For production evaluation, measure more than code syntax.

Useful metrics include:

  • accepted implementation rate;
  • first-pass test success;
  • number of patch iterations;
  • repeated file reads;
  • tool calls;
  • failed commands;
  • unnecessary edits;
  • human intervention;
  • total tokens to acceptance.

This makes 3.5 Flash particularly valuable as a historical baseline for later Gemini Flash generations, which progressively focused on reducing reasoning steps, output tokens, and tool-call overhead.

Rapid coding loops
  1. 01

    Explore

    Inspect repository structure, dependencies, implementation options, and relevant context.

  2. 02

    Implement

    Make coordinated changes across files instead of returning isolated code fragments.

  3. 03

    Execute

    Run code, tests, or tools and use the results as evidence for the next step.

  4. 04

    Iterate

    Repair failures and test alternate paths until the engineering task meets acceptance criteria.

04 / Thought preservation

Automatic Thought Preservation Across Turns

Gemini 3.5 Flash introduced automatic preservation of intermediate reasoning across multi-turn conversations, improving continuity for long debugging, refactoring, and agent workflows.

Reasoning Context Can Carry Forward

In a long agent session, each new turn should not behave like a fresh request that has forgotten how the previous result was reached.

Google’s Gemini 3.5 Flash documentation specifically calls out thought preservation.

With the Interactions API, reasoning state is preserved automatically.

With GenerateContent, reasoning context from previous turns can carry forward when the application passes the full unmodified conversation history, including thought signatures. Google SDKs handle this automatically when used correctly.

This is useful for:

  • iterative debugging;
  • code refactoring;
  • long research tasks;
  • tool-heavy agents;
  • multi-stage business workflows;
  • conversations where an earlier reasoning path still matters several turns later.

Thought preservation can improve continuity, but it can also increase token usage because more reasoning state remains relevant to later requests.

Evaluate whether the quality gain justifies the context cost.

Multi-turn reasoning state
  1. 01

    Preserve

    Carry prior reasoning context forward instead of restarting the decision process on every turn.

  2. 02

    Continue

    Use earlier debugging, planning, and tool observations when choosing the next action.

  3. 03

    Pass history correctly

    For GenerateContent, retain the complete unmodified history and thought signatures when continuing a conversation.

  4. 04

    Measure context growth

    Track whether preserved reasoning materially improves task completion enough to justify additional context usage.

05 / Computer Use

Computer Use Became a Built-In Gemini 3.5 Flash Tool

In June 2026, Google integrated Computer Use directly into Gemini 3.5 Flash, allowing the same model to reason about a task and interact with browser, mobile, and desktop interfaces.

From Tool Calls to Interface Actions

Before this integration, Google exposed Computer Use through a specialized model.

With Gemini 3.5 Flash, Computer Use became a built-in tool in the primary Flash model.

This is important for agents that need to act in systems without a clean API.

A computer-use agent can:

  • inspect the current screen;
  • reason about UI state;
  • click controls;
  • type into fields;
  • navigate between pages;
  • observe the updated interface;
  • continue until the task is complete.

Google positioned this capability for browser automation, enterprise applications, continuous software testing, and knowledge-work workflows.

Computer Use remains a preview capability and requires stronger production safeguards than ordinary text generation.

Evaluation should include wrong clicks, destructive actions, navigation loops, state recovery, confirmation gates, and successful task termination.

Native computer interaction
  1. 01

    See

    Interpret screenshots and current application state.

  2. 02

    Reason

    Choose the next UI action based on the task and observed interface.

  3. 03

    Act

    Click, type, navigate, and interact across supported browser, desktop, or mobile workflows.

  4. 04

    Verify

    Observe the result and continue or recover until the workflow actually reaches its target state.

06 / Pricing

Gemini 3.5 Flash Pricing

Gemini 3.5 Flash costs $1.50 per million input tokens and $9 per million output tokens on the standard paid Gemini Developer API tier.

Thinking Tokens Are Billed as Output

Google’s current standard pricing lists:

  • $1.50 per 1M input tokens;
  • $9.00 per 1M output tokens, including thinking;
  • $0.15 per 1M cached-context tokens;
  • $1.00 per 1M cached tokens per hour of storage.

Batch API pricing is discounted to:

  • $0.75 per 1M input tokens;
  • $4.50 per 1M output tokens.

The Gemini Developer API also lists a free tier for standard 3.5 Flash token usage.

Search and Maps grounding have separate request-based pricing after their shared monthly free allowance.

Because thinking tokens are billed as output, the final visible answer is not enough to estimate cost.

A high-thinking coding task can generate substantial billable reasoning even when the final response is concise.

Standard and batch pricing

1M tokens · USD

Input
$1.50
Cached input
$0.15
Output + thinking
$9.00

Example: 20K input + 4K output

Input cost
$0.0300
Output cost
$0.0360
Estimated total before extra thinking output
$0.0660

07 / Multimodal context

A 1M-Token Native Multimodal Context Window

Gemini 3.5 Flash supports up to 1,048,576 input tokens across text, images, video, audio, and PDFs, with up to 65,536 output tokens.

Long Context Can Hold the Whole Working Set

For an agent, context is not just the user’s prompt.

It can include:

  • repository files;
  • screenshots;
  • long PDFs;
  • video;
  • audio;
  • tool outputs;
  • retrieved URLs;
  • system instructions;
  • prior conversation turns;
  • preserved reasoning state.

Gemini 3.5 Flash can reason across these modalities in one model.

That is useful for long-horizon coding, multimodal research, application automation, document workflows, and agents that need substantial state.

But capacity is not the same as relevance.

Sending every available file, turn, and tool result can increase cost and make important instructions harder to identify.

Use retrieval, caching, selective history, media-resolution controls, and context pruning to keep the working set useful.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Text, images, video, audio, PDFs, conversation history, tool results, and preserved reasoning all contribute to the working context.

08 / Thinking

Minimal, Low, Medium, and High Thinking

Gemini 3.5 Flash supports four thinking levels and defaults to medium, giving applications explicit control over the quality, latency, and cost tradeoff.

Medium Replaced High as the Default

The earlier Gemini 3 Flash Preview defaulted to high thinking.

Gemini 3.5 Flash changed the default to medium, which Google recommends for most workloads because it balances quality, speed, and cost.

Supported levels are:

  • minimal — optimized for response speed and simple requests;
  • low — lower latency and cost while retaining useful reasoning;
  • medium — default and recommended for most code and agent tasks;
  • high — deeper dynamic reasoning for difficult mathematics, coding, and tool orchestration.

Google specifically improved low for code and agentic tasks that need fewer steps.

thinking_budget remains available for backward compatibility, but Google recommends thinking_level. The two parameters cannot be used together.

Thinking levels
minimallowmedium · defaulthigh

Latency-sensitive task

Evaluate minimal or low for chat-like interactions, simple tool calls, and workloads where faster completion matters more than deep reasoning.

Difficult agent task

Use medium or high when stronger planning and tool orchestration improve coding, debugging, mathematics, or long-horizon completion.

09 / Tool ecosystem

Search, Maps, Files, Code, URLs, Functions, and Computer Use

Gemini 3.5 Flash supports a broad tool surface that lets the model move from multimodal reasoning into retrieval, code execution, custom application actions, and graphical-interface interaction.

Combined Tool Use Matters for Agents

Google lists support for:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • URL context;
  • function calling;
  • structured outputs;
  • context caching;
  • Computer Use in preview;
  • Batch API;
  • Flex inference;
  • Priority inference.

Gemini 3 also supports combined tool use.

An agent can, for example, use Search for current evidence, inspect a URL, execute code to calculate a result, call an application-defined function, and return a structured response.

This makes tool orchestration itself an evaluation target.

Measure wrong-tool calls, unnecessary calls, invalid arguments, tool-result interpretation, recovery from failures, and whether the model knows when to stop.

Built-in capabilities
  • Google Search

    Ground responses in current web information.

    Supported
  • Google Maps

    Use Maps grounding for location-aware workflows.

    Supported
  • File Search

    Retrieve relevant information from indexed files.

    Supported
  • Code execution

    Run code during analysis and verification.

    Supported
  • Function calling

    Invoke application-defined tools and business operations.

    Supported
  • Structured outputs

    Return machine-readable responses that follow an expected schema.

    Supported
  • Computer Use

    Preview support for browser, desktop, and mobile interface interaction.

    Supported
  • URL context

    Read and reason over content from supplied URLs.

    Supported

10 / 3.5 → 3.6

Gemini 3.5 Flash vs Gemini 3.6 Flash

Gemini 3.6 Flash keeps the same 1M/64K context profile and broad Gemini tool surface but shifts the emphasis from frontier-speed breakthrough to token-efficient agent execution.

3.6 Targets the Overhead 3.5 Exposed

Gemini 3.5 Flash established the high-capability agentic Flash tier.

Gemini 3.6 Flash then focused on reducing the cost of operating that tier at scale.

Google reported that 3.6 uses roughly 17% fewer output tokens than 3.5 on the Artificial Analysis Index and can reduce output usage much more on some coding evaluations.

Google also describes 3.6 as using fewer reasoning steps and tool calls.

That makes migration a question of workflow efficiency, not basic model capability.

Replay the same agent trajectories and compare:

  • accepted-result rate;
  • output and thinking tokens;
  • tool-call count;
  • repeated actions;
  • latency;
  • human intervention;
  • total cost.

Pricing also changes.

Gemini 3.5 Flash currently costs $1.50 input and $9 output per million tokens.

Gemini 3.6 Flash is temporarily priced at $0.75/$3.75 through December 31, 2026, with $1.50/$7.50 scheduled from January 1, 2027.

Migration comparison
  1. 01

    Keep the same capacity baseline

    Both models use a 1,048,576-token input window and a 65,536-token maximum output.

  2. 02

    Measure reasoning overhead

    Compare thinking and output tokens per successful task rather than only final-answer quality.

  3. 03

    Count tool calls

    3.6 is designed to reduce unnecessary reasoning steps and tool usage, so replay complete agent trajectories.

  4. 04

    Recalculate unit economics

    3.6 has lower output pricing and can also reduce output-token consumption, which compounds savings on agentic workloads.

11 / Evaluation

Gemini 3.5 Flash Strengths and Limitations

Gemini 3.5 Flash remains an important stable baseline for frontier-level Flash agents: strong on coding, long-horizon workflows, multimodality, persistent reasoning, and Computer Use, but less efficient than later Flash generations.

Strengths

  • Frontier-level Flash milestone

    Gemini 3.5 Flash was launched as the first Gemini 3.5 model combining frontier intelligence with Flash-class speed and action.

  • Long-horizon agentic coding

    The model is explicitly designed for rapid coding loops, sub-agent deployment, multi-step workflows, and long-running task execution.

  • Automatic thought preservation

    Reasoning state can carry across multi-turn conversations, improving continuity in debugging, refactoring, and complex agent sessions.

  • Native Computer Use

    Browser, desktop, and mobile interaction is integrated as a built-in preview tool rather than requiring a separate specialist model.

What to consider

  • Higher output price than newer Flash

    The current $9/MTok output rate is above Gemini 3.6, 3.7, and 3.8 Flash pricing.

  • More agent overhead than 3.6

    Google built 3.6 specifically to reduce output tokens, reasoning steps, and tool calls relative to 3.5.

  • Computer Use remains preview

    GUI automation needs permission controls, confirmation gates, auditability, and workload-specific regression tests.

  • Thought preservation can grow context

    Keeping reasoning state improves continuity but can increase token consumption over long conversations.

Benchmark the first frontier Flash generation

Test Gemini 3.5 Flash on your real agent workflows

Replay coding, browser automation, multimodal reasoning, long-horizon tool use, and multi-turn debugging to measure task completion, thinking overhead, tool-call reliability, latency, and total cost.

Start Free

Gemini 3.5 Flash is especially useful as a baseline for measuring whether newer Flash generations reduce token and tool overhead without sacrificing successful task completion.

Common Questions

What is Gemini 3.5 Flash?

Gemini 3.5 Flash is Google’s stable Flash model for frontier-level coding, agentic execution, long-horizon workflows, multimodal reasoning, and fast action.

What is the Gemini 3.5 Flash model ID?

The stable Gemini API model ID is gemini-3.5-flash.

When was Gemini 3.5 Flash released?

Google introduced Gemini 3.5 Flash on May 19, 2026 as the first model in the Gemini 3.5 family.

Is Gemini 3.5 Flash stable?

Yes. Google lists Gemini 3.5 Flash as generally available, stable, and ready for scaled production use.

What is the Gemini 3.5 Flash context window?

Gemini 3.5 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.5 Flash support?

Gemini 3.5 Flash accepts text, images, video, audio, and PDF input and produces text output.

What is the Gemini 3.5 Flash knowledge cutoff?

Google documents January 2025 as the knowledge cutoff for Gemini 3.5 Flash. Use Search grounding for current external information.

How much does Gemini 3.5 Flash cost?

Standard paid Gemini Developer API pricing is $1.50 per 1M input tokens and $9 per 1M output tokens, including thinking tokens.

How much does Gemini 3.5 Flash context caching cost?

Google lists cache reads at $0.15 per 1M cached tokens plus $1 per 1M cached tokens per hour of storage.

Does Gemini 3.5 Flash have a free tier?

Yes. Google currently lists free standard token usage for Gemini 3.5 Flash in the Gemini Developer API, subject to provider limits and product terms.

What are the Gemini 3.5 Flash Batch API prices?

Batch pricing is $0.75 per 1M input tokens and $4.50 per 1M output tokens, including thinking tokens.

Does Gemini 3.5 Flash support thinking?

Yes. It supports minimal, low, medium, and high thinking levels. Medium is the default.

What does minimal thinking do on Gemini 3.5 Flash?

Minimal is optimized for response speed and behaves like a no-thinking mode for most simple requests, although Google notes that the model may still reason minimally on complex tasks.

What is automatic thought preservation?

Gemini 3.5 Flash can carry intermediate reasoning context across multi-turn conversations. This improves continuity for tasks such as iterative debugging and code refactoring when conversation history is preserved correctly.

Does Gemini 3.5 Flash support Computer Use?

Yes, in preview. Google integrated Computer Use directly into Gemini 3.5 Flash in June 2026 for browser, mobile, desktop, testing, and enterprise automation workflows.

What tools does Gemini 3.5 Flash support?

Google lists Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use. Batch, Flex, and Priority inference are also supported.

How is Gemini 3.5 Flash different from Gemini 3.6 Flash?

Both provide a 1M input window, 65,536-token output, multimodal input, and broad tool support. Gemini 3.5 established the frontier-level agentic Flash tier, while Gemini 3.6 focuses more strongly on efficiency by reducing output tokens, reasoning steps, and tool calls.

Should I use Gemini 3.5 Flash for a new application?

It remains a stable production model, but for new workloads you should compare it directly with newer Gemini 3.6, 3.7, and 3.8 Flash models. Keep 3.5 where its quality, latency, Computer Use behavior, or established integrations produce the best cost per accepted result.

Model information

Last updated

Specifications, pricing, stable status, thinking levels, multimodal inputs, automatic thought preservation, Computer Use, tool support, launch positioning, and migration guidance on this page are based on official Google Gemini API, Google DeepMind model-card, and Google product documentation.

Gemini 3.5 Flash — Pricing, 1M Context, Agentic Coding & Computer Use | EidoStack