Google previous-generation Flash model
Gemini 3.6 Flash
A token-efficient Gemini Flash workhorse for coding, knowledge work, multimodal analysis, computer use, and rapid agentic loops—designed to complete real workflows with fewer reasoning steps and tool calls.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $0.75
- intro · per 1M
- Cached input
- $0.075
- intro · per 1M
- Output
- $3.75
- intro · per 1M
01 / Overview
What Gemini 3.6 Flash Is
Gemini 3.6 Flash is Google’s previous-generation stable Flash model for coding, agentic execution, spatial reasoning, multimodal knowledge work, and production workflows where token efficiency matters alongside capability.
A Workhorse Built Around Efficient Agent Execution
Google introduced Gemini 3.6 Flash on July 21, 2026 as the next Flash workhorse after Gemini 3.5 Flash.
Its central design goal was not simply to generate better answers. Google emphasized better quality with fewer output tokens, fewer reasoning steps, and fewer tool calls.
That makes Gemini 3.6 Flash particularly useful as an evaluation baseline for agent economics.
A model can have a low per-token price and still be expensive if it repeatedly calls tools, loops through failed plans, over-generates intermediate reasoning, or needs several retries before the task is accepted.
Gemini 3.6 Flash was built to reduce that overhead.
The model supports a 1,048,576-token input context window and up to 65,536 output tokens. It accepts text, images, video, audio, and PDFs and produces text.
Google currently lists gemini-3.6-flash as a stable model, while its pricing documentation describes it as the previous-generation Flash model.
- Stable model ID:
gemini-3.6-flash. - Introduced July 21, 2026.
- 1,048,576 input tokens.
- 65,536 maximum output tokens.
- Text, image, video, audio, and PDF input.
- Text output.
- Thinking levels: minimal, low, medium, high.
- Default thinking level: medium.
- Provider
- Family
- Gemini 3.6 Flash
- Model ID
- gemini-3.6-flash
- Status
- Stable · previous generation
- Knowledge cutoff
- Mar 2026*
- Input
- Text, Image, Video, Audio, PDF
- Output
- Text
02 / Token efficiency
Fewer Output Tokens, Reasoning Steps, and Tool Calls
Token efficiency is the most distinctive part of Gemini 3.6 Flash’s positioning: Google designed it to complete agentic work with less inference overhead than Gemini 3.5 Flash.
Optimize the Whole Agent Loop
Google reported that Gemini 3.6 Flash used about 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index.
In some software-engineering evaluations, the reduction was much larger. Google cited up to 65% lower output-token usage in DeepSWE.
More importantly for agents, Google states that 3.6 Flash takes fewer reasoning steps and tool calls to complete multi-step workflows.
That changes how the model should be evaluated.
For an agent, useful efficiency metrics include:
- output and thinking tokens per completed task;
- tool calls per successful trajectory;
- duplicated searches or file reads;
- failed execution loops;
- retries after invalid actions;
- total latency to accepted result;
- human interventions;
- cost per completed task.
The cheapest request is not necessarily the cheapest workflow. A model that spends slightly more on one step can still be more efficient if it avoids three corrective steps later.
- 01
Fewer output tokens
Google reported roughly 17% lower output-token usage than Gemini 3.5 Flash on the Artificial Analysis Index.
- 02
Fewer reasoning steps
The model is designed to reach useful conclusions with less unnecessary deliberation across multi-step workflows.
- 03
Fewer tool calls
Reduced tool overhead can lower latency and cost in search, coding, computer-use, and business automation loops.
- 04
Measure completed-task cost
Compare the total trajectory rather than evaluating token price or individual completions in isolation.
03 / Coding & agents
Coding With Fewer Unwanted Edits and Execution Loops
Gemini 3.6 Flash improved software engineering not only by raising capability, but by reducing unnecessary edits and repeated execution cycles.
Precision Matters in Repository-Level Work
A coding agent can fail expensively even when the generated code looks plausible.
It may touch files that do not need changes, overwrite working behavior, rerun the same failing command, misinterpret a test result, or call tools repeatedly without making progress.
Google reported stronger coding performance for 3.6 Flash than 3.5 Flash and highlighted higher precision with fewer unwanted code edits and reduced execution loops.
This makes 3.6 Flash useful for:
- repository maintenance;
- code migrations;
- debugging;
- dependency updates;
- multi-file implementation;
- terminal-based agent workflows;
- iterative test-and-fix loops.
Google also showed Gemini 3.6 Flash performing code migrations through multi-agent orchestration with lower latency and higher quality than 3.5 Flash.
For production evaluation, do not stop at code-generation quality. Measure changed-file precision, test pass rate, number of repair loops, tool execution count, and whether the agent actually finishes the engineering task.
- 01
Inspect
Read the repository, relevant files, errors, and dependencies before changing code.
- 02
Change precisely
Prefer necessary edits over broad rewrites that increase regression risk.
- 03
Execute
Run tests, commands, or code to validate assumptions instead of relying only on generated explanations.
- 04
Avoid loops
Track repeated failed actions and unnecessary tool calls as first-class agent quality metrics.
04 / Knowledge work
Multimodal Knowledge Work Is a Core 3.6 Flash Use Case
Google highlighted Gemini 3.6 Flash improvements in document parsing, chart and data analysis, report drafting, financial information, transcripts, and other knowledge-heavy enterprise workflows.
Combine Understanding With Action
Many enterprise tasks are not pure question answering.
A model may need to read a long PDF, inspect charts, compare figures, analyze a transcript, retrieve current information, run calculations, and then produce a structured report or call an internal tool.
Gemini 3.6 Flash’s multimodal input and built-in tools make these compound workflows possible in one model.
Google cited customer use cases around legal and financial knowledge work and demonstrated 3.6 Flash analyzing financial data and transcripts through Managed Agents.
This makes the model especially relevant for:
- financial document review;
- legal and policy analysis;
- chart interpretation;
- report drafting;
- transcript analysis;
- multimodal research;
- enterprise data extraction.
For evaluation, separate understanding accuracy from workflow accuracy. The model can interpret a document correctly but still choose the wrong downstream action.
Measure extracted facts, cross-document consistency, chart reasoning, citation grounding, tool correctness, structured-output validity, and completion without manual repair.
- 01
Parse
Read PDFs, documents, transcripts, images, and structured business material.
- 02
Reason
Combine facts across text, charts, tables, and multimodal evidence.
- 03
Verify
Use search, code execution, URLs, or tools when the workflow needs external evidence or calculations.
- 04
Deliver
Produce reports, structured outputs, or downstream actions suitable for an application workflow.
05 / Pricing
Gemini 3.6 Flash Pricing
Gemini 3.6 Flash currently uses introductory pricing through December 31, 2026, with higher standard rates scheduled for January 1, 2027.
Current Pricing Is Lower Than the Original Launch Rate
Google launched Gemini 3.6 Flash around a $1.50 input / $7.50 output rate.
The current Gemini Developer API pricing page applies a temporary introductory discount through the end of 2026:
- $0.75 per 1M input tokens;
- $3.75 per 1M output tokens, including thinking tokens;
- $0.075 per 1M cached-context tokens;
- $0.50 per 1M cached tokens per hour of storage.
Starting January 1, 2027, Google lists:
- $1.50 per 1M input tokens;
- $7.50 per 1M output tokens;
- $0.15 per 1M cached-context tokens;
- $1.00 per 1M cached tokens per hour of storage.
Batch and Flex inference provide discounted serving paths, while Priority inference is also supported.
For production budgeting, calculate both current and 2027 economics. An application launched under the temporary discount may face materially different inference costs after the introductory period ends.
1M tokens · USD
- Input
- $0.75
- Cached input
- $0.075
- Output + thinking
- $3.75
Example: 20K input + 4K output
- Input cost
- $0.0150
- Output cost
- $0.0150
- Estimated introductory total
- $0.0300
06 / Multimodal context
A 1M-Token Multimodal Working Context
Gemini 3.6 Flash accepts up to 1,048,576 input tokens across text, images, video, audio, and PDFs and supports up to 65,536 output tokens.
One Context for Code, Documents, Media, and Tool State
The large context window is especially valuable for agent workflows because the working set can contain much more than user messages.
A coding agent can combine source files, logs, architecture documents, screenshots, tool results, and prior steps.
A knowledge-work agent can combine PDFs, charts, audio transcripts, URLs, retrieved evidence, and business instructions.
A computer-use workflow can preserve the task objective alongside screenshots and interaction state.
Native multimodality means these inputs do not need to be split across separate specialist models before reasoning begins.
Still, the full 1M-token ceiling should not be treated as a target. Retrieval, caching, selective history, and context pruning can improve both economics and relevance.
Input context
1,048,576
Max output
65,536
The shared input window can contain text, images, video, audio, PDFs, agent history, instructions, and tool results.
07 / Thinking
Minimal, Low, Medium, and High Thinking
Gemini 3.6 Flash supports four thinking levels—minimal, low, medium, and high—making it more configurable at the low-reasoning end than Gemini 3.7 Flash and Gemini 3.8 Flash.
minimal Is the Distinctive Compatibility Point
The default thinking level is medium.
Use minimal when you want behavior closest to a no-thinking path for simple, high-throughput requests. Google notes that minimal does not guarantee absolutely zero reasoning; the model can still perform very limited thinking for complex tasks.
Use low when latency and cost matter but some reasoning remains useful.
Use medium for balanced everyday production work.
Use high when complex coding, multimodal analysis, or multi-step tool orchestration benefits from deeper deliberation.
This is a meaningful difference from Gemini 3.7 Flash and 3.8 Flash. Those newer models support low, medium, and high, but reject minimal.
Thinking tokens are billed as output tokens, so selecting a higher level affects both latency and cost.
High-throughput task
Evaluate minimal or low for extraction, classification, simple transformations, and latency-sensitive interactions.
Agentic task
Use medium or high where better planning and verification reduce failed loops or incorrect tool calls.
08 / Tools & agents
The Default Model for Google Managed Agents in July 2026
Google made Gemini 3.6 Flash the default model for its Managed Agents preview, reinforcing its role as a balanced model for reasoning, coding, and tool use.
A Model Designed to Operate, Not Only Respond
Managed Agents in the Gemini Interactions API can coordinate reasoning, code execution, package installation, file management, and web retrieval inside an isolated cloud sandbox.
Google switched the Antigravity managed agent to Gemini 3.6 Flash by default on July 28, 2026.
The base model itself supports:
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- URL context;
- function calling;
- structured outputs;
- context caching;
- Computer Use in preview;
- Batch API;
- Flex inference;
- Priority inference.
Computer Use is especially relevant because Google reported improved OSWorld performance for 3.6 Flash compared with 3.5 Flash.
In production, evaluate tool behavior by trajectory: wrong-tool calls, invalid parameters, duplicated calls, recovery after failure, and successful task termination.
- Supported
Google Search
Ground responses in current web information.
- Supported
Google Maps
Use Maps grounding for location-aware workflows.
- Supported
File Search
Retrieve information from indexed files.
- Supported
Code execution
Execute code as part of analysis, validation, and agent workflows.
- Supported
Function calling
Invoke application-defined tools and actions.
- Supported
Structured outputs
Return schema-constrained machine-readable responses.
- Supported
Computer Use
Preview support for interacting with graphical interfaces.
- Supported
URL context
Read and reason over supplied web URLs.
09 / Safety
Enhanced Frontier Safety Safeguards
Gemini 3.6 Flash shipped with strengthened safeguards for frontier-risk domains while Google also worked to reduce unnecessary refusals for beneficial use cases.
Safety Behavior Is Part of Production Reliability
Google specifically highlighted enhanced protections for chemical, biological, radiological, and nuclear risk and cyber-offense misuse.
The company also reported increased resistance to jailbreaks.
For production applications, safety should be included in normal model evaluation rather than treated as a separate compliance exercise.
A useful test set should include:
- clearly benign requests;
- ambiguous requests;
- requests close to policy boundaries;
- legitimate security or scientific use cases;
- prompt-injection and jailbreak attempts;
- downstream tool actions triggered by risky content.
The goal is to measure both sides of the behavior: whether unsafe requests are handled appropriately and whether legitimate workflows avoid unnecessary refusal.
This is particularly relevant for agent systems because a model can move from text generation into tools, code execution, files, or external actions.
- 01
CBRN protections
Google strengthened mitigations around chemical, biological, radiological, and nuclear misuse.
- 02
Cyber safeguards
The model includes stronger protections against cyber-offense misuse.
- 03
Jailbreak resistance
Google reports substantially improved resistance to attempts to bypass safety controls.
- 04
Benign-use evaluation
Test refusal rates on legitimate workflows so stronger safety does not silently reduce useful completion rates.
10 / 3.6 → 3.7
Gemini 3.6 Flash vs Gemini 3.7 Flash
Gemini 3.7 Flash keeps the same 1M/64K capacity and the same 2026 introductory token pricing, but shifts the Flash baseline toward stronger first-pass coding, web development, knowledge work, and more diligent multi-step execution.
The Most Important API Difference Is minimal
Both models support:
- 1,048,576 input tokens;
- 65,536 output tokens;
- text, image, video, audio, and PDF input;
- multimodal reasoning;
- the broad Gemini tool ecosystem;
- the same introductory and scheduled standard token rates.
The key low-level thinking difference is that Gemini 3.6 Flash supports minimal, while Gemini 3.7 Flash does not.
An application that sends thinking_level: minimal must migrate that path to low.
Behaviorally, Google positioned 3.7 around stronger production coding, first-pass accuracy, web-development quality, knowledge work, and more deliberate agent planning.
That means 3.6 can remain a useful efficiency baseline.
If 3.6 already completes a workload reliably with fewer reasoning tokens or with minimal thinking, upgrading may not improve the economics. If coding tasks require repeated repairs or agent workflows need stronger planning, 3.7 can justify the migration.
- 01
Replace minimal with low
Gemini 3.7 Flash rejects minimal thinking, so migrate low-reasoning traffic to low and re-run latency and quality evaluations.
- 02
Keep capacity assumptions
Both models use the same 1,048,576-token input and 65,536-token output limits.
- 03
Replay coding failures
Use tasks where 3.6 required extra patches or tool loops to measure whether 3.7 improves first-pass completion.
- 04
Compare cost per accepted result
Do not upgrade only because the version number is newer; measure reasoning tokens, retries, tool calls, latency, and accepted-task rate.
11 / Evaluation
Gemini 3.6 Flash Strengths and Limitations
Gemini 3.6 Flash remains a useful stable efficiency baseline for coding, multimodal knowledge work, and agent loops, even though newer Flash models now provide stronger execution on difficult workloads.
Strengths
Token-efficient agent execution
Google designed 3.6 Flash to reduce output tokens, reasoning steps, and tool calls compared with 3.5 Flash.
Minimal thinking support
The model can use minimal reasoning for latency-sensitive or high-throughput work, a setting removed from 3.7 and 3.8 Flash.
Strong coding and computer use
3.6 improved code precision, execution-loop behavior, and computer-use performance over its predecessor.
Broad multimodal knowledge work
A 1M-token context and native text, image, video, audio, and PDF input support document-heavy and enterprise workflows.
What to consider
Previous-generation Flash
Google now describes 3.6 Flash as the previous-generation Flash model, with 3.7 and 3.8 offering newer capability improvements.
Introductory pricing expires
Current $0.75/$3.75 pricing is scheduled to become $1.50/$7.50 on January 1, 2027.
Newer models improve difficult trajectories
Gemini 3.7 and especially 3.8 move further toward first-pass coding quality, long-horizon software engineering, and autonomous agent reliability.
Foundation-model limitations remain
Google’s model card notes hallucinations as well as occasional slowness or timeout issues, so production validation remains necessary.
Measure efficiency at the workflow level
Test Gemini 3.6 Flash on your real agent workload
Compare coding, document analysis, computer use, tool-heavy loops, multimodal prompts, and long-context tasks to measure accepted-result quality, reasoning overhead, tool-call count, latency, and total cost per completed workflow.
Start FreeGemini 3.6 Flash is most interesting when fewer reasoning steps and tool calls translate into lower end-to-end cost—not merely when its per-token price looks attractive.
Common Questions
What is Gemini 3.6 Flash?
Gemini 3.6 Flash is Google’s stable previous-generation Flash model for coding, agentic execution, spatial reasoning, multimodal knowledge work, and real-world workflows where token efficiency matters.
What is the Gemini 3.6 Flash model ID?
The stable Gemini API model ID is gemini-3.6-flash.
When was Gemini 3.6 Flash released?
Google introduced Gemini 3.6 Flash on July 21, 2026 and made it available through the Gemini API and other Google AI products.
Is Gemini 3.6 Flash deprecated?
No. Google currently lists gemini-3.6-flash as a stable model. Its pricing page describes it as the previous-generation Flash model, but no shutdown date is currently listed.
What is the Gemini 3.6 Flash context window?
Gemini 3.6 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3.6 Flash support?
The model accepts text, images, video, audio, and PDF input and generates text output.
What is the Gemini 3.6 Flash knowledge cutoff?
Google DeepMind lists March 2026 as the Gemini 3.6 Flash knowledge cutoff, while noting that some domains may still be limited to January 2025 in line with the broader Gemini 3 family.
How much does Gemini 3.6 Flash cost?
Through December 31, 2026, standard paid pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, Google lists $1.50 input and $7.50 output per million tokens.
How much does Gemini 3.6 Flash context caching cost?
Through December 31, 2026, cached context costs $0.075 per 1M tokens plus $0.50 per 1M tokens per hour of storage. Starting January 1, 2027, Google lists $0.15 per 1M cached tokens plus $1.00 per 1M tokens per hour.
Does Gemini 3.6 Flash support thinking?
Yes. Gemini 3.6 Flash supports minimal, low, medium, and high thinking levels. Medium is the default.
What does minimal thinking mean on Gemini 3.6 Flash?
Minimal is the lowest supported thinking level and is intended to approximate a no-thinking path for simple or high-throughput requests. Google notes that it does not guarantee absolutely zero reasoning on every task.
Does Gemini 3.7 Flash support the same minimal thinking level?
No. Gemini 3.7 Flash and Gemini 3.8 Flash support low, medium, and high but reject minimal. Applications migrating from 3.6 should map minimal workloads to low and re-evaluate latency and output quality.
Is Gemini 3.6 Flash good for coding?
Yes. Google positions code generation and agentic coding loops as core use cases and reported improved precision, fewer unwanted edits, and reduced execution loops compared with Gemini 3.5 Flash.
Why is Gemini 3.6 Flash considered token efficient?
Google reported roughly 17% lower output-token consumption than Gemini 3.5 Flash on the Artificial Analysis Index and stated that 3.6 Flash uses fewer reasoning steps and tool calls for multi-step workflows.
Does Gemini 3.6 Flash support Computer Use?
Yes, in preview. Google also reported improved computer-use benchmark performance compared with Gemini 3.5 Flash.
What tools does Gemini 3.6 Flash support?
The official model page lists Google Search grounding, Google Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use.
Was Gemini 3.6 Flash used for Google Managed Agents?
Yes. Google made Gemini 3.6 Flash the default model for its Antigravity Managed Agents preview in July 2026, describing it as a balanced model for reasoning, coding, and tool use.
How is Gemini 3.6 Flash different from Gemini 3.7 Flash?
Both use the same 1M input and 65,536 output limits and current pricing structure. Gemini 3.6 focuses strongly on token-efficient agent execution and supports minimal thinking, while Gemini 3.7 moves toward stronger first-pass coding, web development, knowledge work, and more deliberate multi-step planning.
Should I use Gemini 3.6 Flash for a new application?
Evaluate it when token-efficient agent execution or minimal-thinking latency is valuable. For new coding and long-horizon agent workloads, also test Gemini 3.7 and 3.8 Flash and choose based on cost per accepted result rather than model generation alone.
Model information
Last updated
Specifications, pricing, model status, thinking levels, multimodal inputs, token-efficiency claims, Managed Agents usage, safety positioning, tool support, and migration guidance on this page are based on official Google Gemini API documentation, Google DeepMind’s Gemini 3.6 Flash model card, and Google product announcements.