Google Gemini stable model
Gemini 3.5 Flash
The first Gemini 3.5 Flash model to combine frontier-level intelligence with fast agentic action for coding, long-horizon workflows, multimodal reasoning, and computer-use automation.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $1.50
- per 1M tokens
- Cached input
- $0.15
- per 1M tokens
- Output
- $9.00
- per 1M tokens
01 / Overview
What Gemini 3.5 Flash Is
Gemini 3.5 Flash is Google’s stable Flash model that brought frontier-level coding, agentic execution, and long-horizon reasoning into the high-speed Gemini Flash tier.
The First Gemini 3.5 Model
Google introduced Gemini 3.5 Flash on May 19, 2026 as the first model in the Gemini 3.5 family.
The launch theme was frontier intelligence with action.
That positioning matters because Gemini 3.5 Flash was not introduced as a lightweight fallback for simple requests. Google explicitly targeted difficult coding, autonomous agents, multi-step workflows, sub-agent deployment, and long-horizon tasks.
Google also reported that the model outperformed Gemini 3.1 Pro on several coding and agentic benchmarks while running substantially faster than other frontier models.
The stable Gemini API model ID is gemini-3.5-flash.
It accepts text, images, video, audio, and PDFs and generates text. The model supports up to 1,048,576 input tokens and up to 65,536 output tokens.
Gemini 3.5 Flash also supports dynamic thinking, automatic thought preservation, Google’s built-in tool ecosystem, context caching, and Computer Use in preview.
- Stable model ID:
gemini-3.5-flash. - Released May 19, 2026.
- Stable and GA for scaled production use.
- 1,048,576 input tokens.
- 65,536 maximum output tokens.
- Text, image, video, audio, and PDF input.
- Text output.
- Default thinking level: medium.
- Provider
- Family
- Gemini 3.5 Flash
- Model ID
- gemini-3.5-flash
- Status
- Stable · GA
- Knowledge cutoff
- Jan 2025
- Input
- Text, Image, Video, Audio, PDF
- Output
- Text
02 / Frontier + action
Frontier Intelligence at Flash Speed
Gemini 3.5 Flash was designed around a specific product idea: combine high-end reasoning capability with the responsiveness and throughput traditionally associated with Flash.
Quality and Latency Are Both Part of the Model
At launch, Google described 3.5 Flash as delivering intelligence comparable with large flagship models while preserving Flash-class speed.
Google also reported roughly four-times-higher output-token throughput than other frontier models in its launch comparisons.
That makes the model relevant to applications where a strong result is not enough if it arrives too slowly.
Examples include:
- interactive coding agents;
- browser automation;
- real-time operational analysis;
- multi-step customer workflows;
- document-heavy productivity tools;
- agent orchestration where one model call triggers several downstream actions.
For these systems, evaluate both task quality and time to useful completion.
A model can have excellent benchmark scores but still be the wrong product choice if long reasoning delays every user-visible action.
Gemini 3.5 Flash’s SEO and evaluation identity therefore sits between two older assumptions: it is more capable than a traditional speed-first model, but less expensive and faster than a Pro-tier frontier model.
- 01
Reason
Handle difficult coding, multimodal, and multi-step problems with frontier-level capability.
- 02
Act
Move from reasoning into tools, sub-agents, browser interaction, code execution, and application-defined functions.
- 03
Iterate quickly
Use Flash-class responsiveness for workflows that require repeated reason-act-observe loops.
- 04
Measure completion latency
Track total time to a successful outcome instead of comparing isolated response latency.
03 / Agentic coding
Built for Rapid Agentic Coding Loops
Google describes Gemini 3.5 Flash as particularly effective for rapid agentic loops involving complex coding cycles, iterative exploration, and long-running software tasks.
Coding Quality Is a Trajectory Property
Repository-level engineering rarely succeeds in one generation.
A coding agent may need to:
- inspect source files;
- understand dependencies;
- form a plan;
- modify several files;
- execute tests;
- diagnose failures;
- search for more context;
- repair the implementation;
- verify the final state.
Gemini 3.5 Flash was explicitly built for this loop.
Google’s launch materials emphasized coding, sub-agent deployment, long-horizon task execution, and rapid iteration over alternate approaches.
For production evaluation, measure more than code syntax.
Useful metrics include:
- accepted implementation rate;
- first-pass test success;
- number of patch iterations;
- repeated file reads;
- tool calls;
- failed commands;
- unnecessary edits;
- human intervention;
- total tokens to acceptance.
This makes 3.5 Flash particularly valuable as a historical baseline for later Gemini Flash generations, which progressively focused on reducing reasoning steps, output tokens, and tool-call overhead.
- 01
Explore
Inspect repository structure, dependencies, implementation options, and relevant context.
- 02
Implement
Make coordinated changes across files instead of returning isolated code fragments.
- 03
Execute
Run code, tests, or tools and use the results as evidence for the next step.
- 04
Iterate
Repair failures and test alternate paths until the engineering task meets acceptance criteria.
04 / Thought preservation
Automatic Thought Preservation Across Turns
Gemini 3.5 Flash introduced automatic preservation of intermediate reasoning across multi-turn conversations, improving continuity for long debugging, refactoring, and agent workflows.
Reasoning Context Can Carry Forward
In a long agent session, each new turn should not behave like a fresh request that has forgotten how the previous result was reached.
Google’s Gemini 3.5 Flash documentation specifically calls out thought preservation.
With the Interactions API, reasoning state is preserved automatically.
With GenerateContent, reasoning context from previous turns can carry forward when the application passes the full unmodified conversation history, including thought signatures. Google SDKs handle this automatically when used correctly.
This is useful for:
- iterative debugging;
- code refactoring;
- long research tasks;
- tool-heavy agents;
- multi-stage business workflows;
- conversations where an earlier reasoning path still matters several turns later.
Thought preservation can improve continuity, but it can also increase token usage because more reasoning state remains relevant to later requests.
Evaluate whether the quality gain justifies the context cost.
- 01
Preserve
Carry prior reasoning context forward instead of restarting the decision process on every turn.
- 02
Continue
Use earlier debugging, planning, and tool observations when choosing the next action.
- 03
Pass history correctly
For GenerateContent, retain the complete unmodified history and thought signatures when continuing a conversation.
- 04
Measure context growth
Track whether preserved reasoning materially improves task completion enough to justify additional context usage.
05 / Computer Use
Computer Use Became a Built-In Gemini 3.5 Flash Tool
In June 2026, Google integrated Computer Use directly into Gemini 3.5 Flash, allowing the same model to reason about a task and interact with browser, mobile, and desktop interfaces.
From Tool Calls to Interface Actions
Before this integration, Google exposed Computer Use through a specialized model.
With Gemini 3.5 Flash, Computer Use became a built-in tool in the primary Flash model.
This is important for agents that need to act in systems without a clean API.
A computer-use agent can:
- inspect the current screen;
- reason about UI state;
- click controls;
- type into fields;
- navigate between pages;
- observe the updated interface;
- continue until the task is complete.
Google positioned this capability for browser automation, enterprise applications, continuous software testing, and knowledge-work workflows.
Computer Use remains a preview capability and requires stronger production safeguards than ordinary text generation.
Evaluation should include wrong clicks, destructive actions, navigation loops, state recovery, confirmation gates, and successful task termination.
- 01
See
Interpret screenshots and current application state.
- 02
Reason
Choose the next UI action based on the task and observed interface.
- 03
Act
Click, type, navigate, and interact across supported browser, desktop, or mobile workflows.
- 04
Verify
Observe the result and continue or recover until the workflow actually reaches its target state.
06 / Pricing
Gemini 3.5 Flash Pricing
Gemini 3.5 Flash costs $1.50 per million input tokens and $9 per million output tokens on the standard paid Gemini Developer API tier.
Thinking Tokens Are Billed as Output
Google’s current standard pricing lists:
- $1.50 per 1M input tokens;
- $9.00 per 1M output tokens, including thinking;
- $0.15 per 1M cached-context tokens;
- $1.00 per 1M cached tokens per hour of storage.
Batch API pricing is discounted to:
- $0.75 per 1M input tokens;
- $4.50 per 1M output tokens.
The Gemini Developer API also lists a free tier for standard 3.5 Flash token usage.
Search and Maps grounding have separate request-based pricing after their shared monthly free allowance.
Because thinking tokens are billed as output, the final visible answer is not enough to estimate cost.
A high-thinking coding task can generate substantial billable reasoning even when the final response is concise.
1M tokens · USD
- Input
- $1.50
- Cached input
- $0.15
- Output + thinking
- $9.00
Example: 20K input + 4K output
- Input cost
- $0.0300
- Output cost
- $0.0360
- Estimated total before extra thinking output
- $0.0660
07 / Multimodal context
A 1M-Token Native Multimodal Context Window
Gemini 3.5 Flash supports up to 1,048,576 input tokens across text, images, video, audio, and PDFs, with up to 65,536 output tokens.
Long Context Can Hold the Whole Working Set
For an agent, context is not just the user’s prompt.
It can include:
- repository files;
- screenshots;
- long PDFs;
- video;
- audio;
- tool outputs;
- retrieved URLs;
- system instructions;
- prior conversation turns;
- preserved reasoning state.
Gemini 3.5 Flash can reason across these modalities in one model.
That is useful for long-horizon coding, multimodal research, application automation, document workflows, and agents that need substantial state.
But capacity is not the same as relevance.
Sending every available file, turn, and tool result can increase cost and make important instructions harder to identify.
Use retrieval, caching, selective history, media-resolution controls, and context pruning to keep the working set useful.
Input context
1,048,576
Max output
65,536
Text, images, video, audio, PDFs, conversation history, tool results, and preserved reasoning all contribute to the working context.
08 / Thinking
Minimal, Low, Medium, and High Thinking
Gemini 3.5 Flash supports four thinking levels and defaults to medium, giving applications explicit control over the quality, latency, and cost tradeoff.
Medium Replaced High as the Default
The earlier Gemini 3 Flash Preview defaulted to high thinking.
Gemini 3.5 Flash changed the default to medium, which Google recommends for most workloads because it balances quality, speed, and cost.
Supported levels are:
minimal— optimized for response speed and simple requests;low— lower latency and cost while retaining useful reasoning;medium— default and recommended for most code and agent tasks;high— deeper dynamic reasoning for difficult mathematics, coding, and tool orchestration.
Google specifically improved low for code and agentic tasks that need fewer steps.
thinking_budget remains available for backward compatibility, but Google recommends thinking_level. The two parameters cannot be used together.
Latency-sensitive task
Evaluate minimal or low for chat-like interactions, simple tool calls, and workloads where faster completion matters more than deep reasoning.
Difficult agent task
Use medium or high when stronger planning and tool orchestration improve coding, debugging, mathematics, or long-horizon completion.
09 / Tool ecosystem
Search, Maps, Files, Code, URLs, Functions, and Computer Use
Gemini 3.5 Flash supports a broad tool surface that lets the model move from multimodal reasoning into retrieval, code execution, custom application actions, and graphical-interface interaction.
Combined Tool Use Matters for Agents
Google lists support for:
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- URL context;
- function calling;
- structured outputs;
- context caching;
- Computer Use in preview;
- Batch API;
- Flex inference;
- Priority inference.
Gemini 3 also supports combined tool use.
An agent can, for example, use Search for current evidence, inspect a URL, execute code to calculate a result, call an application-defined function, and return a structured response.
This makes tool orchestration itself an evaluation target.
Measure wrong-tool calls, unnecessary calls, invalid arguments, tool-result interpretation, recovery from failures, and whether the model knows when to stop.
- Supported
Google Search
Ground responses in current web information.
- Supported
Google Maps
Use Maps grounding for location-aware workflows.
- Supported
File Search
Retrieve relevant information from indexed files.
- Supported
Code execution
Run code during analysis and verification.
- Supported
Function calling
Invoke application-defined tools and business operations.
- Supported
Structured outputs
Return machine-readable responses that follow an expected schema.
- Supported
Computer Use
Preview support for browser, desktop, and mobile interface interaction.
- Supported
URL context
Read and reason over content from supplied URLs.
10 / 3.5 → 3.6
Gemini 3.5 Flash vs Gemini 3.6 Flash
Gemini 3.6 Flash keeps the same 1M/64K context profile and broad Gemini tool surface but shifts the emphasis from frontier-speed breakthrough to token-efficient agent execution.
3.6 Targets the Overhead 3.5 Exposed
Gemini 3.5 Flash established the high-capability agentic Flash tier.
Gemini 3.6 Flash then focused on reducing the cost of operating that tier at scale.
Google reported that 3.6 uses roughly 17% fewer output tokens than 3.5 on the Artificial Analysis Index and can reduce output usage much more on some coding evaluations.
Google also describes 3.6 as using fewer reasoning steps and tool calls.
That makes migration a question of workflow efficiency, not basic model capability.
Replay the same agent trajectories and compare:
- accepted-result rate;
- output and thinking tokens;
- tool-call count;
- repeated actions;
- latency;
- human intervention;
- total cost.
Pricing also changes.
Gemini 3.5 Flash currently costs $1.50 input and $9 output per million tokens.
Gemini 3.6 Flash is temporarily priced at $0.75/$3.75 through December 31, 2026, with $1.50/$7.50 scheduled from January 1, 2027.
- 01
Keep the same capacity baseline
Both models use a 1,048,576-token input window and a 65,536-token maximum output.
- 02
Measure reasoning overhead
Compare thinking and output tokens per successful task rather than only final-answer quality.
- 03
Count tool calls
3.6 is designed to reduce unnecessary reasoning steps and tool usage, so replay complete agent trajectories.
- 04
Recalculate unit economics
3.6 has lower output pricing and can also reduce output-token consumption, which compounds savings on agentic workloads.
11 / Evaluation
Gemini 3.5 Flash Strengths and Limitations
Gemini 3.5 Flash remains an important stable baseline for frontier-level Flash agents: strong on coding, long-horizon workflows, multimodality, persistent reasoning, and Computer Use, but less efficient than later Flash generations.
Strengths
Frontier-level Flash milestone
Gemini 3.5 Flash was launched as the first Gemini 3.5 model combining frontier intelligence with Flash-class speed and action.
Long-horizon agentic coding
The model is explicitly designed for rapid coding loops, sub-agent deployment, multi-step workflows, and long-running task execution.
Automatic thought preservation
Reasoning state can carry across multi-turn conversations, improving continuity in debugging, refactoring, and complex agent sessions.
Native Computer Use
Browser, desktop, and mobile interaction is integrated as a built-in preview tool rather than requiring a separate specialist model.
What to consider
Higher output price than newer Flash
The current $9/MTok output rate is above Gemini 3.6, 3.7, and 3.8 Flash pricing.
More agent overhead than 3.6
Google built 3.6 specifically to reduce output tokens, reasoning steps, and tool calls relative to 3.5.
Computer Use remains preview
GUI automation needs permission controls, confirmation gates, auditability, and workload-specific regression tests.
Thought preservation can grow context
Keeping reasoning state improves continuity but can increase token consumption over long conversations.
Benchmark the first frontier Flash generation
Test Gemini 3.5 Flash on your real agent workflows
Replay coding, browser automation, multimodal reasoning, long-horizon tool use, and multi-turn debugging to measure task completion, thinking overhead, tool-call reliability, latency, and total cost.
Start FreeGemini 3.5 Flash is especially useful as a baseline for measuring whether newer Flash generations reduce token and tool overhead without sacrificing successful task completion.
Common Questions
What is Gemini 3.5 Flash?
Gemini 3.5 Flash is Google’s stable Flash model for frontier-level coding, agentic execution, long-horizon workflows, multimodal reasoning, and fast action.
What is the Gemini 3.5 Flash model ID?
The stable Gemini API model ID is gemini-3.5-flash.
When was Gemini 3.5 Flash released?
Google introduced Gemini 3.5 Flash on May 19, 2026 as the first model in the Gemini 3.5 family.
Is Gemini 3.5 Flash stable?
Yes. Google lists Gemini 3.5 Flash as generally available, stable, and ready for scaled production use.
What is the Gemini 3.5 Flash context window?
Gemini 3.5 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3.5 Flash support?
Gemini 3.5 Flash accepts text, images, video, audio, and PDF input and produces text output.
What is the Gemini 3.5 Flash knowledge cutoff?
Google documents January 2025 as the knowledge cutoff for Gemini 3.5 Flash. Use Search grounding for current external information.
How much does Gemini 3.5 Flash cost?
Standard paid Gemini Developer API pricing is $1.50 per 1M input tokens and $9 per 1M output tokens, including thinking tokens.
How much does Gemini 3.5 Flash context caching cost?
Google lists cache reads at $0.15 per 1M cached tokens plus $1 per 1M cached tokens per hour of storage.
Does Gemini 3.5 Flash have a free tier?
Yes. Google currently lists free standard token usage for Gemini 3.5 Flash in the Gemini Developer API, subject to provider limits and product terms.
What are the Gemini 3.5 Flash Batch API prices?
Batch pricing is $0.75 per 1M input tokens and $4.50 per 1M output tokens, including thinking tokens.
Does Gemini 3.5 Flash support thinking?
Yes. It supports minimal, low, medium, and high thinking levels. Medium is the default.
What does minimal thinking do on Gemini 3.5 Flash?
Minimal is optimized for response speed and behaves like a no-thinking mode for most simple requests, although Google notes that the model may still reason minimally on complex tasks.
What is automatic thought preservation?
Gemini 3.5 Flash can carry intermediate reasoning context across multi-turn conversations. This improves continuity for tasks such as iterative debugging and code refactoring when conversation history is preserved correctly.
Does Gemini 3.5 Flash support Computer Use?
Yes, in preview. Google integrated Computer Use directly into Gemini 3.5 Flash in June 2026 for browser, mobile, desktop, testing, and enterprise automation workflows.
What tools does Gemini 3.5 Flash support?
Google lists Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use. Batch, Flex, and Priority inference are also supported.
How is Gemini 3.5 Flash different from Gemini 3.6 Flash?
Both provide a 1M input window, 65,536-token output, multimodal input, and broad tool support. Gemini 3.5 established the frontier-level agentic Flash tier, while Gemini 3.6 focuses more strongly on efficiency by reducing output tokens, reasoning steps, and tool calls.
Should I use Gemini 3.5 Flash for a new application?
It remains a stable production model, but for new workloads you should compare it directly with newer Gemini 3.6, 3.7, and 3.8 Flash models. Keep 3.5 where its quality, latency, Computer Use behavior, or established integrations produce the best cost per accepted result.
Model information
Last updated
Specifications, pricing, stable status, thinking levels, multimodal inputs, automatic thought preservation, Computer Use, tool support, launch positioning, and migration guidance on this page are based on official Google Gemini API, Google DeepMind model-card, and Google product documentation.