Google previous-generation Flash model

Gemini 3.6 Flash

A token-efficient Gemini Flash workhorse for coding, knowledge work, multimodal analysis, computer use, and rapid agentic loops—designed to complete real workflows with fewer reasoning steps and tool calls.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.75
intro · per 1M
Cached input
$0.075
intro · per 1M
Output
$3.75
intro · per 1M

01 / Overview

What Gemini 3.6 Flash Is

Gemini 3.6 Flash is Google’s previous-generation stable Flash model for coding, agentic execution, spatial reasoning, multimodal knowledge work, and production workflows where token efficiency matters alongside capability.

A Workhorse Built Around Efficient Agent Execution

Google introduced Gemini 3.6 Flash on July 21, 2026 as the next Flash workhorse after Gemini 3.5 Flash.

Its central design goal was not simply to generate better answers. Google emphasized better quality with fewer output tokens, fewer reasoning steps, and fewer tool calls.

That makes Gemini 3.6 Flash particularly useful as an evaluation baseline for agent economics.

A model can have a low per-token price and still be expensive if it repeatedly calls tools, loops through failed plans, over-generates intermediate reasoning, or needs several retries before the task is accepted.

Gemini 3.6 Flash was built to reduce that overhead.

The model supports a 1,048,576-token input context window and up to 65,536 output tokens. It accepts text, images, video, audio, and PDFs and produces text.

Google currently lists gemini-3.6-flash as a stable model, while its pricing documentation describes it as the previous-generation Flash model.

  • Stable model ID: gemini-3.6-flash.
  • Introduced July 21, 2026.
  • 1,048,576 input tokens.
  • 65,536 maximum output tokens.
  • Text, image, video, audio, and PDF input.
  • Text output.
  • Thinking levels: minimal, low, medium, high.
  • Default thinking level: medium.
Model profile
Provider
Google
Family
Gemini 3.6 Flash
Model ID
gemini-3.6-flash
Status
Stable · previous generation
Knowledge cutoff
Mar 2026*
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / Token efficiency

Fewer Output Tokens, Reasoning Steps, and Tool Calls

Token efficiency is the most distinctive part of Gemini 3.6 Flash’s positioning: Google designed it to complete agentic work with less inference overhead than Gemini 3.5 Flash.

Optimize the Whole Agent Loop

Google reported that Gemini 3.6 Flash used about 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index.

In some software-engineering evaluations, the reduction was much larger. Google cited up to 65% lower output-token usage in DeepSWE.

More importantly for agents, Google states that 3.6 Flash takes fewer reasoning steps and tool calls to complete multi-step workflows.

That changes how the model should be evaluated.

For an agent, useful efficiency metrics include:

  • output and thinking tokens per completed task;
  • tool calls per successful trajectory;
  • duplicated searches or file reads;
  • failed execution loops;
  • retries after invalid actions;
  • total latency to accepted result;
  • human interventions;
  • cost per completed task.

The cheapest request is not necessarily the cheapest workflow. A model that spends slightly more on one step can still be more efficient if it avoids three corrective steps later.

Efficiency improvements
  1. 01

    Fewer output tokens

    Google reported roughly 17% lower output-token usage than Gemini 3.5 Flash on the Artificial Analysis Index.

  2. 02

    Fewer reasoning steps

    The model is designed to reach useful conclusions with less unnecessary deliberation across multi-step workflows.

  3. 03

    Fewer tool calls

    Reduced tool overhead can lower latency and cost in search, coding, computer-use, and business automation loops.

  4. 04

    Measure completed-task cost

    Compare the total trajectory rather than evaluating token price or individual completions in isolation.

03 / Coding & agents

Coding With Fewer Unwanted Edits and Execution Loops

Gemini 3.6 Flash improved software engineering not only by raising capability, but by reducing unnecessary edits and repeated execution cycles.

Precision Matters in Repository-Level Work

A coding agent can fail expensively even when the generated code looks plausible.

It may touch files that do not need changes, overwrite working behavior, rerun the same failing command, misinterpret a test result, or call tools repeatedly without making progress.

Google reported stronger coding performance for 3.6 Flash than 3.5 Flash and highlighted higher precision with fewer unwanted code edits and reduced execution loops.

This makes 3.6 Flash useful for:

  • repository maintenance;
  • code migrations;
  • debugging;
  • dependency updates;
  • multi-file implementation;
  • terminal-based agent workflows;
  • iterative test-and-fix loops.

Google also showed Gemini 3.6 Flash performing code migrations through multi-agent orchestration with lower latency and higher quality than 3.5 Flash.

For production evaluation, do not stop at code-generation quality. Measure changed-file precision, test pass rate, number of repair loops, tool execution count, and whether the agent actually finishes the engineering task.

Efficient coding loops
  1. 01

    Inspect

    Read the repository, relevant files, errors, and dependencies before changing code.

  2. 02

    Change precisely

    Prefer necessary edits over broad rewrites that increase regression risk.

  3. 03

    Execute

    Run tests, commands, or code to validate assumptions instead of relying only on generated explanations.

  4. 04

    Avoid loops

    Track repeated failed actions and unnecessary tool calls as first-class agent quality metrics.

04 / Knowledge work

Multimodal Knowledge Work Is a Core 3.6 Flash Use Case

Google highlighted Gemini 3.6 Flash improvements in document parsing, chart and data analysis, report drafting, financial information, transcripts, and other knowledge-heavy enterprise workflows.

Combine Understanding With Action

Many enterprise tasks are not pure question answering.

A model may need to read a long PDF, inspect charts, compare figures, analyze a transcript, retrieve current information, run calculations, and then produce a structured report or call an internal tool.

Gemini 3.6 Flash’s multimodal input and built-in tools make these compound workflows possible in one model.

Google cited customer use cases around legal and financial knowledge work and demonstrated 3.6 Flash analyzing financial data and transcripts through Managed Agents.

This makes the model especially relevant for:

  • financial document review;
  • legal and policy analysis;
  • chart interpretation;
  • report drafting;
  • transcript analysis;
  • multimodal research;
  • enterprise data extraction.

For evaluation, separate understanding accuracy from workflow accuracy. The model can interpret a document correctly but still choose the wrong downstream action.

Measure extracted facts, cross-document consistency, chart reasoning, citation grounding, tool correctness, structured-output validity, and completion without manual repair.

Documents, charts, and reports
  1. 01

    Parse

    Read PDFs, documents, transcripts, images, and structured business material.

  2. 02

    Reason

    Combine facts across text, charts, tables, and multimodal evidence.

  3. 03

    Verify

    Use search, code execution, URLs, or tools when the workflow needs external evidence or calculations.

  4. 04

    Deliver

    Produce reports, structured outputs, or downstream actions suitable for an application workflow.

05 / Pricing

Gemini 3.6 Flash Pricing

Gemini 3.6 Flash currently uses introductory pricing through December 31, 2026, with higher standard rates scheduled for January 1, 2027.

Current Pricing Is Lower Than the Original Launch Rate

Google launched Gemini 3.6 Flash around a $1.50 input / $7.50 output rate.

The current Gemini Developer API pricing page applies a temporary introductory discount through the end of 2026:

  • $0.75 per 1M input tokens;
  • $3.75 per 1M output tokens, including thinking tokens;
  • $0.075 per 1M cached-context tokens;
  • $0.50 per 1M cached tokens per hour of storage.

Starting January 1, 2027, Google lists:

  • $1.50 per 1M input tokens;
  • $7.50 per 1M output tokens;
  • $0.15 per 1M cached-context tokens;
  • $1.00 per 1M cached tokens per hour of storage.

Batch and Flex inference provide discounted serving paths, while Priority inference is also supported.

For production budgeting, calculate both current and 2027 economics. An application launched under the temporary discount may face materially different inference costs after the introductory period ends.

Introductory and standard pricing

1M tokens · USD

Input
$0.75
Cached input
$0.075
Output + thinking
$3.75

Example: 20K input + 4K output

Input cost
$0.0150
Output cost
$0.0150
Estimated introductory total
$0.0300

06 / Multimodal context

A 1M-Token Multimodal Working Context

Gemini 3.6 Flash accepts up to 1,048,576 input tokens across text, images, video, audio, and PDFs and supports up to 65,536 output tokens.

One Context for Code, Documents, Media, and Tool State

The large context window is especially valuable for agent workflows because the working set can contain much more than user messages.

A coding agent can combine source files, logs, architecture documents, screenshots, tool results, and prior steps.

A knowledge-work agent can combine PDFs, charts, audio transcripts, URLs, retrieved evidence, and business instructions.

A computer-use workflow can preserve the task objective alongside screenshots and interaction state.

Native multimodality means these inputs do not need to be split across separate specialist models before reasoning begins.

Still, the full 1M-token ceiling should not be treated as a target. Retrieval, caching, selective history, and context pruning can improve both economics and relevance.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

The shared input window can contain text, images, video, audio, PDFs, agent history, instructions, and tool results.

07 / Thinking

Minimal, Low, Medium, and High Thinking

Gemini 3.6 Flash supports four thinking levels—minimal, low, medium, and high—making it more configurable at the low-reasoning end than Gemini 3.7 Flash and Gemini 3.8 Flash.

minimal Is the Distinctive Compatibility Point

The default thinking level is medium.

Use minimal when you want behavior closest to a no-thinking path for simple, high-throughput requests. Google notes that minimal does not guarantee absolutely zero reasoning; the model can still perform very limited thinking for complex tasks.

Use low when latency and cost matter but some reasoning remains useful.

Use medium for balanced everyday production work.

Use high when complex coding, multimodal analysis, or multi-step tool orchestration benefits from deeper deliberation.

This is a meaningful difference from Gemini 3.7 Flash and 3.8 Flash. Those newer models support low, medium, and high, but reject minimal.

Thinking tokens are billed as output tokens, so selecting a higher level affects both latency and cost.

Thinking levels
minimal · supportedlowmedium · defaulthigh

High-throughput task

Evaluate minimal or low for extraction, classification, simple transformations, and latency-sensitive interactions.

Agentic task

Use medium or high where better planning and verification reduce failed loops or incorrect tool calls.

08 / Tools & agents

The Default Model for Google Managed Agents in July 2026

Google made Gemini 3.6 Flash the default model for its Managed Agents preview, reinforcing its role as a balanced model for reasoning, coding, and tool use.

A Model Designed to Operate, Not Only Respond

Managed Agents in the Gemini Interactions API can coordinate reasoning, code execution, package installation, file management, and web retrieval inside an isolated cloud sandbox.

Google switched the Antigravity managed agent to Gemini 3.6 Flash by default on July 28, 2026.

The base model itself supports:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • URL context;
  • function calling;
  • structured outputs;
  • context caching;
  • Computer Use in preview;
  • Batch API;
  • Flex inference;
  • Priority inference.

Computer Use is especially relevant because Google reported improved OSWorld performance for 3.6 Flash compared with 3.5 Flash.

In production, evaluate tool behavior by trajectory: wrong-tool calls, invalid parameters, duplicated calls, recovery after failure, and successful task termination.

Built-in agent surface
  • Google Search

    Ground responses in current web information.

    Supported
  • Google Maps

    Use Maps grounding for location-aware workflows.

    Supported
  • File Search

    Retrieve information from indexed files.

    Supported
  • Code execution

    Execute code as part of analysis, validation, and agent workflows.

    Supported
  • Function calling

    Invoke application-defined tools and actions.

    Supported
  • Structured outputs

    Return schema-constrained machine-readable responses.

    Supported
  • Computer Use

    Preview support for interacting with graphical interfaces.

    Supported
  • URL context

    Read and reason over supplied web URLs.

    Supported

09 / Safety

Enhanced Frontier Safety Safeguards

Gemini 3.6 Flash shipped with strengthened safeguards for frontier-risk domains while Google also worked to reduce unnecessary refusals for beneficial use cases.

Safety Behavior Is Part of Production Reliability

Google specifically highlighted enhanced protections for chemical, biological, radiological, and nuclear risk and cyber-offense misuse.

The company also reported increased resistance to jailbreaks.

For production applications, safety should be included in normal model evaluation rather than treated as a separate compliance exercise.

A useful test set should include:

  • clearly benign requests;
  • ambiguous requests;
  • requests close to policy boundaries;
  • legitimate security or scientific use cases;
  • prompt-injection and jailbreak attempts;
  • downstream tool actions triggered by risky content.

The goal is to measure both sides of the behavior: whether unsafe requests are handled appropriately and whether legitimate workflows avoid unnecessary refusal.

This is particularly relevant for agent systems because a model can move from text generation into tools, code execution, files, or external actions.

Frontier safeguards
  1. 01

    CBRN protections

    Google strengthened mitigations around chemical, biological, radiological, and nuclear misuse.

  2. 02

    Cyber safeguards

    The model includes stronger protections against cyber-offense misuse.

  3. 03

    Jailbreak resistance

    Google reports substantially improved resistance to attempts to bypass safety controls.

  4. 04

    Benign-use evaluation

    Test refusal rates on legitimate workflows so stronger safety does not silently reduce useful completion rates.

10 / 3.6 → 3.7

Gemini 3.6 Flash vs Gemini 3.7 Flash

Gemini 3.7 Flash keeps the same 1M/64K capacity and the same 2026 introductory token pricing, but shifts the Flash baseline toward stronger first-pass coding, web development, knowledge work, and more diligent multi-step execution.

The Most Important API Difference Is minimal

Both models support:

  • 1,048,576 input tokens;
  • 65,536 output tokens;
  • text, image, video, audio, and PDF input;
  • multimodal reasoning;
  • the broad Gemini tool ecosystem;
  • the same introductory and scheduled standard token rates.

The key low-level thinking difference is that Gemini 3.6 Flash supports minimal, while Gemini 3.7 Flash does not.

An application that sends thinking_level: minimal must migrate that path to low.

Behaviorally, Google positioned 3.7 around stronger production coding, first-pass accuracy, web-development quality, knowledge work, and more deliberate agent planning.

That means 3.6 can remain a useful efficiency baseline.

If 3.6 already completes a workload reliably with fewer reasoning tokens or with minimal thinking, upgrading may not improve the economics. If coding tasks require repeated repairs or agent workflows need stronger planning, 3.7 can justify the migration.

Migration comparison
  1. 01

    Replace minimal with low

    Gemini 3.7 Flash rejects minimal thinking, so migrate low-reasoning traffic to low and re-run latency and quality evaluations.

  2. 02

    Keep capacity assumptions

    Both models use the same 1,048,576-token input and 65,536-token output limits.

  3. 03

    Replay coding failures

    Use tasks where 3.6 required extra patches or tool loops to measure whether 3.7 improves first-pass completion.

  4. 04

    Compare cost per accepted result

    Do not upgrade only because the version number is newer; measure reasoning tokens, retries, tool calls, latency, and accepted-task rate.

11 / Evaluation

Gemini 3.6 Flash Strengths and Limitations

Gemini 3.6 Flash remains a useful stable efficiency baseline for coding, multimodal knowledge work, and agent loops, even though newer Flash models now provide stronger execution on difficult workloads.

Strengths

  • Token-efficient agent execution

    Google designed 3.6 Flash to reduce output tokens, reasoning steps, and tool calls compared with 3.5 Flash.

  • Minimal thinking support

    The model can use minimal reasoning for latency-sensitive or high-throughput work, a setting removed from 3.7 and 3.8 Flash.

  • Strong coding and computer use

    3.6 improved code precision, execution-loop behavior, and computer-use performance over its predecessor.

  • Broad multimodal knowledge work

    A 1M-token context and native text, image, video, audio, and PDF input support document-heavy and enterprise workflows.

What to consider

  • Previous-generation Flash

    Google now describes 3.6 Flash as the previous-generation Flash model, with 3.7 and 3.8 offering newer capability improvements.

  • Introductory pricing expires

    Current $0.75/$3.75 pricing is scheduled to become $1.50/$7.50 on January 1, 2027.

  • Newer models improve difficult trajectories

    Gemini 3.7 and especially 3.8 move further toward first-pass coding quality, long-horizon software engineering, and autonomous agent reliability.

  • Foundation-model limitations remain

    Google’s model card notes hallucinations as well as occasional slowness or timeout issues, so production validation remains necessary.

Measure efficiency at the workflow level

Test Gemini 3.6 Flash on your real agent workload

Compare coding, document analysis, computer use, tool-heavy loops, multimodal prompts, and long-context tasks to measure accepted-result quality, reasoning overhead, tool-call count, latency, and total cost per completed workflow.

Start Free

Gemini 3.6 Flash is most interesting when fewer reasoning steps and tool calls translate into lower end-to-end cost—not merely when its per-token price looks attractive.

Common Questions

What is Gemini 3.6 Flash?

Gemini 3.6 Flash is Google’s stable previous-generation Flash model for coding, agentic execution, spatial reasoning, multimodal knowledge work, and real-world workflows where token efficiency matters.

What is the Gemini 3.6 Flash model ID?

The stable Gemini API model ID is gemini-3.6-flash.

When was Gemini 3.6 Flash released?

Google introduced Gemini 3.6 Flash on July 21, 2026 and made it available through the Gemini API and other Google AI products.

Is Gemini 3.6 Flash deprecated?

No. Google currently lists gemini-3.6-flash as a stable model. Its pricing page describes it as the previous-generation Flash model, but no shutdown date is currently listed.

What is the Gemini 3.6 Flash context window?

Gemini 3.6 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.6 Flash support?

The model accepts text, images, video, audio, and PDF input and generates text output.

What is the Gemini 3.6 Flash knowledge cutoff?

Google DeepMind lists March 2026 as the Gemini 3.6 Flash knowledge cutoff, while noting that some domains may still be limited to January 2025 in line with the broader Gemini 3 family.

How much does Gemini 3.6 Flash cost?

Through December 31, 2026, standard paid pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, Google lists $1.50 input and $7.50 output per million tokens.

How much does Gemini 3.6 Flash context caching cost?

Through December 31, 2026, cached context costs $0.075 per 1M tokens plus $0.50 per 1M tokens per hour of storage. Starting January 1, 2027, Google lists $0.15 per 1M cached tokens plus $1.00 per 1M tokens per hour.

Does Gemini 3.6 Flash support thinking?

Yes. Gemini 3.6 Flash supports minimal, low, medium, and high thinking levels. Medium is the default.

What does minimal thinking mean on Gemini 3.6 Flash?

Minimal is the lowest supported thinking level and is intended to approximate a no-thinking path for simple or high-throughput requests. Google notes that it does not guarantee absolutely zero reasoning on every task.

Does Gemini 3.7 Flash support the same minimal thinking level?

No. Gemini 3.7 Flash and Gemini 3.8 Flash support low, medium, and high but reject minimal. Applications migrating from 3.6 should map minimal workloads to low and re-evaluate latency and output quality.

Is Gemini 3.6 Flash good for coding?

Yes. Google positions code generation and agentic coding loops as core use cases and reported improved precision, fewer unwanted edits, and reduced execution loops compared with Gemini 3.5 Flash.

Why is Gemini 3.6 Flash considered token efficient?

Google reported roughly 17% lower output-token consumption than Gemini 3.5 Flash on the Artificial Analysis Index and stated that 3.6 Flash uses fewer reasoning steps and tool calls for multi-step workflows.

Does Gemini 3.6 Flash support Computer Use?

Yes, in preview. Google also reported improved computer-use benchmark performance compared with Gemini 3.5 Flash.

What tools does Gemini 3.6 Flash support?

The official model page lists Google Search grounding, Google Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use.

Was Gemini 3.6 Flash used for Google Managed Agents?

Yes. Google made Gemini 3.6 Flash the default model for its Antigravity Managed Agents preview in July 2026, describing it as a balanced model for reasoning, coding, and tool use.

How is Gemini 3.6 Flash different from Gemini 3.7 Flash?

Both use the same 1M input and 65,536 output limits and current pricing structure. Gemini 3.6 focuses strongly on token-efficient agent execution and supports minimal thinking, while Gemini 3.7 moves toward stronger first-pass coding, web development, knowledge work, and more deliberate multi-step planning.

Should I use Gemini 3.6 Flash for a new application?

Evaluate it when token-efficient agent execution or minimal-thinking latency is valuable. For new coding and long-horizon agent workloads, also test Gemini 3.7 and 3.8 Flash and choose based on cost per accepted result rather than model generation alone.

Model information

Last updated

Specifications, pricing, model status, thinking levels, multimodal inputs, token-efficiency claims, Managed Agents usage, safety positioning, tool support, and migration guidance on this page are based on official Google Gemini API documentation, Google DeepMind’s Gemini 3.6 Flash model card, and Google product announcements.

Gemini 3.6 Flash — Pricing, 1M Context, Token Efficiency & Agentic Coding | EidoStack