Google Gemini legacy preview

Gemini 3 Flash Preview

The original Gemini 3 Flash preview that brought Pro-grade reasoning, multimodal intelligence, agentic coding, and rich tool use into a faster and lower-cost Flash model.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.50
text/image/video · per 1M
Cached input
$0.05
text/image/video · per 1M
Output
$3.00
per 1M tokens

01 / Overview

What Gemini 3 Flash Preview Is

Gemini 3 Flash Preview is the original Flash model of the Gemini 3 generation, launched to combine frontier-level intelligence with the latency, efficiency, and price profile expected from Google’s Flash tier.

The First Gemini 3 Flash

Google released gemini-3-flash-preview on December 17, 2025.

At launch, Google described the model as frontier intelligence built for speed.

The important product shift was that Flash was no longer positioned only as a cheaper model for straightforward requests. Gemini 3 Flash brought strong reasoning, multimodal understanding, coding, and agentic behavior into the faster tier.

Google reported that the model surpassed Gemini 2.5 Pro across many benchmarks while operating at lower cost and higher speed.

Today, Google’s pricing documentation describes Gemini 3 Flash Preview as a legacy Flash model providing baseline speed and intelligence.

The endpoint remains available in public preview and currently has no announced shutdown date, but Google’s deprecation table recommends gemini-3.6-flash as its replacement.

The model supports:

  • 1,048,576 input tokens;
  • 65,536 output tokens;
  • text, image, audio, video, and document input;
  • text output;
  • dynamic thinking;
  • structured outputs;
  • context caching;
  • search grounding;
  • code execution;
  • function calling;
  • URL context;
  • Computer Use in preview.
Model profile
Provider
Google
Family
Gemini 3 Flash
Model ID
gemini-3-flash-preview
Launch stage
Public preview · legacy
Release date
Dec 17, 2025
Knowledge cutoff
Jan 2025
Output
Text

02 / Frontier + speed

Pro-Grade Reasoning Without Pro-Tier Latency

Gemini 3 Flash Preview launched around a specific promise: preserve much of Gemini 3’s high-end reasoning while reducing cost and latency enough for interactive and high-frequency applications.

Speed Became Part of Frontier Capability

Before Gemini 3 Flash, developers often treated “frontier reasoning” and “fast production model” as separate categories.

Google explicitly challenged that split.

Gemini 3 Flash was presented as combining Gemini 3’s Pro-grade reasoning foundation with Flash-level latency, efficiency, and cost.

This made it relevant to workloads such as:

  • interactive analysis;
  • coding assistants;
  • rapid agent loops;
  • multimodal applications;
  • document analysis;
  • search-grounded answers;
  • UI generation;
  • tool-heavy application workflows.

Google also made Gemini 3 Flash the default model in the Gemini app and AI Mode in Search at launch, demonstrating that the model was intended for high-frequency user-facing interactions rather than only offline API work.

The correct evaluation metric is therefore not raw intelligence alone.

Measure:

  • accepted-result quality;
  • time to first useful answer;
  • complete-task latency;
  • reasoning-token usage;
  • retries;
  • tool calls;
  • cost per accepted result.
Frontier intelligence at Flash speed
  1. 01

    Reason

    Apply Gemini 3 reasoning to complex questions, coding, multimodal analysis, and agentic workflows.

  2. 02

    Respond quickly

    Use Flash-class latency for interactive experiences where a strong answer must also arrive quickly.

  3. 03

    Act

    Move from reasoning into search, code execution, functions, URLs, and Computer Use.

  4. 04

    Measure the full result

    Evaluate quality, latency, tokens, retries, and tools together rather than optimizing a single benchmark.

03 / Agentic coding

The Gemini 3 Flash Preview Was Built for Agentic Coding

Agentic coding was one of the defining launch workloads for Gemini 3 Flash Preview, with Google positioning the model for high-frequency terminal workflows and iterative software-engineering tasks.

Coding Is More Than Generating a Function

Google brought Gemini 3 Flash into Gemini CLI immediately at launch.

The model was designed for workflows where the AI repeatedly:

  • reads code;
  • searches files;
  • reasons about dependencies;
  • edits implementations;
  • runs commands;
  • interprets failures;
  • retries;
  • validates the final result.

Google reported a 78% SWE-bench Verified score for Gemini 3 Flash in its December 2025 Gemini CLI announcement.

That benchmark number is useful context, but production evaluation should focus on complete engineering trajectories.

Measure:

  • first-pass test success;
  • number of patch iterations;
  • repeated file reads;
  • invalid tool calls;
  • command failures;
  • unwanted edits;
  • human intervention;
  • latency to accepted implementation.

This page therefore treats Gemini 3 Flash Preview as an agentic-coding baseline, not merely as a generic chat model.

Coding and terminal workflows
  1. 01

    Inspect

    Read repository files, logs, errors, dependencies, and surrounding project context.

  2. 02

    Plan

    Choose an implementation path and decide which tools or files are actually needed.

  3. 03

    Execute

    Edit code, invoke tools, run commands, and use execution results as evidence.

  4. 04

    Repair

    Iterate after failed tests or incorrect assumptions until the task meets acceptance criteria.

04 / Visual reasoning

Code Execution Can Turn Vision Into an Active Investigation

One of Gemini 3 Flash Preview’s most distinctive API features is the ability to combine image reasoning with code execution so the model can zoom, crop, annotate, count, or otherwise manipulate visual input during analysis.

The Model Can Inspect an Image Step by Step

A normal vision model receives an image and produces an interpretation.

Gemini 3 Flash Preview can go further.

Google documents a workflow in which the model:

  1. forms a plan;
  2. writes Python;
  3. executes code against the image;
  4. zooms or crops relevant regions;
  5. visually grounds the answer using the manipulated result.

This is useful for:

  • dense diagrams;
  • screenshots;
  • small visual details;
  • object counting;
  • interface inspection;
  • charts;
  • spatial reasoning;
  • technical images.

Google described Gemini 3 Flash at launch as having its most advanced visual and spatial reasoning for the model class.

The important evaluation question is whether active inspection improves accuracy enough to justify extra execution steps and latency.

Active visual investigation
  1. 01

    Observe

    Interpret the initial image and identify which details need closer inspection.

  2. 02

    Plan

    Decide whether cropping, zooming, annotation, counting, or another transformation is necessary.

  3. 03

    Execute

    Write and run code that transforms or inspects the visual input.

  4. 04

    Ground

    Use the manipulated visual evidence to produce a more precise final answer.

05 / Pricing

Gemini 3 Flash Preview Pricing

Gemini 3 Flash Preview uses low Flash pricing: $0.50 per million text, image, or video input tokens and $3 per million output tokens on the standard paid tier.

Audio Input Has a Separate Rate

Google currently lists Standard pricing at:

  • $0.50 per 1M text, image, or video input tokens;
  • $1.00 per 1M audio input tokens;
  • $3.00 per 1M output tokens, including thinking;
  • $0.05 per 1M cached text, image, or video tokens;
  • $0.10 per 1M cached audio tokens;
  • $1.00 per 1M cached tokens per hour of storage.

Batch and Flex pricing reduce token rates to:

  • $0.25 text/image/video input;
  • $0.50 audio input;
  • $1.50 output.

Priority pricing increases rates to:

  • $0.90 text/image/video input;
  • $1.80 audio input;
  • $5.40 output.

The Gemini Developer API also provides a free tier for Gemini 3 Flash Preview standard token usage.

Thinking tokens are included in output billing, so high-default reasoning can make the effective cost larger than the visible final response suggests.

Standard and discounted pricing

1M tokens · USD

Text / image / video input
$0.50
Audio input
$1.00
Cached text / image / video
$0.05
Output + thinking
$3.00

Example: 20K text input + 4K output

Input cost
$0.0100
Output cost
$0.0120
Estimated total before extra thinking
$0.0220

06 / Multimodal context

1,048,576 Tokens of Multimodal Input

Gemini 3 Flash Preview supports a 1,048,576-token context window and up to 65,536 output tokens across text-centric and multimodal workloads.

The Working Context Can Include Much More Than Text

The model can process:

  • text;
  • images;
  • audio;
  • video;
  • PDFs and other documents;
  • conversation history;
  • tool results;
  • system instructions;
  • retrieved URLs.

Google Cloud documents support for up to 3,000 images per prompt, up to 3,000 document pages per file, long-form video, and audio inputs that can extend to roughly 8.4 hours within token limits.

This makes the model useful for:

  • repository analysis;
  • document review;
  • long video understanding;
  • audio summarization;
  • multimodal research;
  • screenshot-heavy workflows;
  • persistent agent sessions.

Large context should still be curated.

Sending irrelevant history or media increases cost and can reduce the model’s ability to focus on the actual task.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Text, images, audio, video, documents, tool results, history, and instructions all share the same working context.

07 / Thinking

High Thinking Is the Default

Gemini 3 Flash Preview supports minimal, low, medium, and high thinking levels, but unlike later stable Flash models it defaults to high.

The Preview Was Tuned Toward Reasoning Depth

Google’s Gemini 3 developer guide documents:

  • minimal;
  • low;
  • medium;
  • high.

high is the default and uses dynamic thinking.

That means an application that does not explicitly configure reasoning may spend more time and output tokens thinking than a later Flash integration that defaults to medium.

Use:

  • minimal for very simple, latency-sensitive requests;
  • low for straightforward instruction following and high-throughput work;
  • medium for a balanced quality/latency tradeoff;
  • high for difficult coding, multimodal reasoning, mathematics, and tool-heavy workflows.

Google retains the older thinking_budget parameter for compatibility but recommends thinking_level. The two cannot be used in the same request.

Thinking levels
minimallowmediumhigh · default

Interactive workload

Reduce thinking to minimal or low when response latency matters and deeper reasoning does not improve acceptance.

Complex agent workload

Keep medium or high for difficult coding, visual reasoning, and multi-step tool orchestration.

08 / Tools

Built-In Tools and Custom Functions Can Work Together

Gemini 3 introduced richer tool composition, allowing built-in Google tools and developer-defined functions to participate in the same workflow.

One Agent Can Retrieve, Compute, and Act

The Gemini 3 developer documentation lists support for:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • URL context;
  • function calling;
  • structured outputs;
  • context caching.

Gemini 3 also supports combining built-in tools with custom function calls.

For example, a workflow can use Google Search to identify current information and then call an application-defined function to act on that information.

Structured outputs can be combined with supported tools so the final grounded answer can still conform to a machine-readable schema.

This makes Gemini 3 Flash Preview useful for agents that must move between retrieval, reasoning, computation, and application actions.

Gemini 3 tool composition
  • Google Search

    Ground responses in current web information beyond the model’s January 2025 knowledge cutoff.

    Supported
  • Google Maps

    Use Maps grounding in supported Gemini 3 workflows.

    Supported
  • File Search

    Retrieve relevant content from indexed file collections.

    Supported
  • Code execution

    Run code for calculations, verification, and active visual investigation.

    Supported
  • Function calling

    Invoke application-defined tools and combine them with supported built-in tools.

    Supported
  • Structured outputs

    Return schema-constrained machine-readable results.

    Supported
  • URL context

    Read and reason over content from supplied URLs.

    Supported
  • Live API

    Google Cloud lists Gemini Live API as unsupported for this model.

    Not listed

09 / Computer Use

Computer Use Was Added to Gemini 3 Flash Preview

Google added Computer Use support to Gemini 3 Flash Preview in January 2026, allowing the model to reason over graphical interfaces and act through browser or UI automation workflows.

The Flash Model Can Work Through Interfaces, Not Only APIs

Computer Use is relevant when the system being automated does not expose the required operation through a clean API.

An agent can inspect visual state, decide the next action, interact with controls, observe the result, and continue.

Potential workloads include:

  • browser automation;
  • application testing;
  • repetitive enterprise workflows;
  • data entry;
  • cross-system operations;
  • UI-based research.

The feature remains in preview.

Production systems should evaluate more than successful demos.

Measure incorrect clicks, navigation loops, destructive actions, recovery after unexpected UI state, confirmation requirements, and successful termination.

GUI automation
  1. 01

    Observe

    Interpret screenshots and current application state.

  2. 02

    Decide

    Reason about the next interaction needed to advance the task.

  3. 03

    Act

    Interact with supported interfaces through the Computer Use tool.

  4. 04

    Verify

    Inspect the updated state and continue, recover, or stop when the target outcome is reached.

10 / Preview lifecycle

Gemini 3 Flash Preview Is Now a Legacy Preview Model

The endpoint remains available, but Google now labels it as a legacy Flash model and recommends moving production workloads to a newer stable Flash release.

No Shutdown Date Is Announced Yet

Google’s deprecation table currently lists:

  • release date: December 17, 2025;
  • status: preview model;
  • shutdown date: no shutdown date announced;
  • recommended replacement: gemini-3.6-flash.

Google’s general model documentation notes that preview models may be used in production but can have more restrictive rate limits and are normally deprecated with advance notice.

That means the model can still be useful for:

  • regression testing;
  • compatibility evaluation;
  • historical model comparisons;
  • applications that have not yet migrated.

It is less attractive as the default choice for a brand-new production system.

Keep the model ID configurable and maintain regression tests so moving to a stable replacement does not require rewriting provider-specific application logic.

Legacy preview status
  1. 01

    Still available

    Google currently lists no shutdown date for gemini-3-flash-preview.

  2. 02

    Legacy classification

    Current pricing documentation describes it as a legacy Flash model providing baseline speed and intelligence.

  3. 03

    Stable replacement

    Google’s current deprecation table recommends gemini-3.6-flash as the replacement.

  4. 04

    Keep migration cheap

    Store model routing and reasoning settings in configuration so a future retirement does not require application-wide code changes.

11 / Migration

Migrating From Gemini 3 Flash Preview

Moving from Gemini 3 Flash Preview to a stable Flash model is not only a model-ID change: reasoning defaults, request configuration, price, and agent behavior should all be re-baselined.

The First Stable Successors Changed the Operating Profile

Google’s Gemini 3.5 Flash migration guide explicitly documents the move from:

gemini-3-flash-preview → gemini-3.5-flash

The guide calls out several changes:

  • default thinking moved from high to medium;
  • thinking_level is preferred over thinking_budget;
  • older sampling overrides such as temperature, top_p, and top_k should be removed from the migration path;
  • function responses need correct IDs and names;
  • prompts should be re-tested for quality, speed, and cost.

Google’s current deprecation table now recommends gemini-3.6-flash as the replacement for the preview endpoint.

That means new migrations should evaluate the current stable Flash family rather than assuming 3.5 is still the best destination.

The economics also differ.

Gemini 3 Flash Preview costs $0.50 input / $3 output for standard text, image, and video usage.

Later stable Flash models can have different prices and reasoning behavior, so migration should compare complete tasks rather than only API compatibility.

Migration to stable Flash
  1. 01

    Update the model target

    Move away from the legacy preview to an evaluated stable Flash endpoint; Google currently recommends Gemini 3.6 Flash.

  2. 02

    Re-test thinking

    The preview defaults to high, while later Flash releases can default to medium or use different supported thinking levels.

  3. 03

    Clean request parameters

    Prefer thinking_level, remove stale sampling overrides where migration guidance recommends it, and validate function-response IDs.

  4. 04

    Re-baseline task economics

    Compare accepted-result rate, tokens, retries, tool calls, latency, and total cost before switching production traffic.

12 / Evaluation

Gemini 3 Flash Preview Strengths and Limitations

Gemini 3 Flash Preview remains an important Gemini 3 baseline: it introduced frontier-grade Flash reasoning, rich multimodal tooling, visual code execution, and Computer Use, but newer stable Flash generations are now the better default for most production deployments.

Strengths

  • Frontier reasoning at Flash economics

    The model established the Gemini 3 pattern of bringing stronger reasoning and multimodal capability into a faster, lower-cost tier.

  • High-default reasoning

    Dynamic high thinking makes the preview useful for comparing reasoning-heavy behavior against later medium-default Flash models.

  • Active visual reasoning

    Code execution can zoom, crop, annotate, count, and manipulate images as part of step-by-step visual investigation.

  • Rich agent tool composition

    Search, Maps, files, code, URLs, functions, structured outputs, caching, and Computer Use support complex multimodal workflows.

What to consider

  • Legacy preview status

    Google now labels the model as legacy and recommends migrating to a stable Flash model.

  • High reasoning by default

    Unconfigured requests can spend more time and output tokens reasoning than later Flash models that default to medium.

  • Preview lifecycle risk

    No shutdown date is currently announced, but preview endpoints can be deprecated and should not be hard-coded deeply into application architecture.

  • Newer Flash generations improve production behavior

    Gemini 3.5 through 3.8 add stronger coding, efficiency, long-horizon execution, and newer agent capabilities, making them more relevant for new deployments.

Benchmark the original Gemini 3 Flash baseline

Test Gemini 3 Flash Preview against stable Flash models

Replay coding, visual reasoning, multimodal analysis, tool use, browser automation, and long-context prompts to compare quality, latency, thinking cost, tool-call reliability, and migration impact.

Start Free

Gemini 3 Flash Preview is now a legacy preview model. Use it as an evaluation or compatibility baseline, and keep production routing ready to move to a stable Gemini Flash endpoint.

Common Questions

What is Gemini 3 Flash Preview?

Gemini 3 Flash Preview is the original Gemini 3 Flash model, launched as a fast and cost-effective model combining frontier-level reasoning, multimodal intelligence, coding, and agentic capabilities.

What is the Gemini 3 Flash Preview model ID?

The model ID is gemini-3-flash-preview.

When was Gemini 3 Flash Preview released?

Google released Gemini 3 Flash Preview on December 17, 2025.

Is Gemini 3 Flash Preview still available?

Yes. Google currently lists the model as a public-preview endpoint with no announced shutdown date, although current pricing documentation describes it as a legacy Flash model.

What model should replace Gemini 3 Flash Preview?

Google’s current Gemini deprecations table recommends gemini-3.6-flash as the replacement. Applications should also evaluate newer stable Flash releases against their own workload requirements.

What is the Gemini 3 Flash Preview context window?

Gemini 3 Flash Preview supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3 Flash Preview support?

The model supports text, images, audio, video, and document input and generates text output.

What is the Gemini 3 Flash Preview knowledge cutoff?

Google documents January 2025 as the knowledge cutoff for Gemini 3 models. Search grounding can be used when current external information is required.

How much does Gemini 3 Flash Preview cost?

Standard paid pricing is $0.50 per 1M text, image, or video input tokens, $1 per 1M audio input tokens, and $3 per 1M output tokens including thinking.

How much does Gemini 3 Flash Preview context caching cost?

Standard cache reads cost $0.05 per 1M text, image, or video tokens and $0.10 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.

Does Gemini 3 Flash Preview have a free tier?

Yes. Google currently lists free standard token usage for Gemini 3 Flash Preview in the Gemini Developer API, subject to provider limits and terms.

What are the Batch prices for Gemini 3 Flash Preview?

Batch pricing is $0.25 per 1M text, image, or video input tokens, $0.50 per 1M audio input tokens, and $1.50 per 1M output tokens.

Does Gemini 3 Flash Preview support thinking?

Yes. It supports minimal, low, medium, and high thinking levels, with high as the default dynamic setting.

Does Gemini 3 Flash Preview support minimal thinking?

Yes. Minimal is supported and is intended to minimize latency for simple chat and high-throughput tasks, although Google notes that the model may still perform very limited reasoning on complex requests.

Can I use thinking_budget with Gemini 3 Flash Preview?

Yes for backward compatibility, but Google recommends thinking_level for more predictable behavior. Do not send thinking_budget and thinking_level together.

Is Gemini 3 Flash Preview good for coding?

Yes. Agentic coding was one of its launch workloads, and Google brought the model to Gemini CLI for high-frequency terminal workflows.

What is code execution with images?

Gemini 3 Flash Preview can reason about an image, write and execute code to crop, zoom, annotate, count, or otherwise manipulate the visual input, then use the transformed evidence to answer more precisely.

Does Gemini 3 Flash Preview support Computer Use?

Yes, in preview. Google added Computer Use support to gemini-3-flash-preview in January 2026.

What tools does Gemini 3 Flash Preview support?

Gemini 3 documentation lists Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, and context caching. Computer Use is also supported in preview.

Does Gemini 3 Flash Preview support the Live API?

No. Google Cloud currently lists Gemini Live API as unsupported for this model.

How is Gemini 3 Flash Preview different from Gemini 3.5 Flash?

Gemini 3 Flash Preview is the original legacy preview with high thinking by default and $0.50/$3 standard text pricing. Gemini 3.5 Flash is a stable successor focused more heavily on frontier agentic coding, automatic thought preservation, and production use, with medium thinking by default and higher token pricing.

Should I use Gemini 3 Flash Preview for a new production application?

For a new production application, a stable Gemini Flash model is generally the better default. Gemini 3 Flash Preview remains useful for compatibility, regression testing, historical comparison, or workloads where its specific price and behavior still outperform the migration alternatives.

Model information

Last updated

Specifications, pricing, launch date, preview status, thinking levels, visual reasoning, code execution with images, multimodal function responses, tools, Computer Use, deprecation guidance, and migration behavior on this page are based on official Google Gemini API, Google Cloud, and Google product documentation.

Gemini 3 Flash Preview — Pricing, 1M Context, High Thinking & Multimodal Tools | EidoStack