Google Gemini legacy preview
Gemini 3 Flash Preview
The original Gemini 3 Flash preview that brought Pro-grade reasoning, multimodal intelligence, agentic coding, and rich tool use into a faster and lower-cost Flash model.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $0.50
- text/image/video · per 1M
- Cached input
- $0.05
- text/image/video · per 1M
- Output
- $3.00
- per 1M tokens
01 / Overview
What Gemini 3 Flash Preview Is
Gemini 3 Flash Preview is the original Flash model of the Gemini 3 generation, launched to combine frontier-level intelligence with the latency, efficiency, and price profile expected from Google’s Flash tier.
The First Gemini 3 Flash
Google released gemini-3-flash-preview on December 17, 2025.
At launch, Google described the model as frontier intelligence built for speed.
The important product shift was that Flash was no longer positioned only as a cheaper model for straightforward requests. Gemini 3 Flash brought strong reasoning, multimodal understanding, coding, and agentic behavior into the faster tier.
Google reported that the model surpassed Gemini 2.5 Pro across many benchmarks while operating at lower cost and higher speed.
Today, Google’s pricing documentation describes Gemini 3 Flash Preview as a legacy Flash model providing baseline speed and intelligence.
The endpoint remains available in public preview and currently has no announced shutdown date, but Google’s deprecation table recommends gemini-3.6-flash as its replacement.
The model supports:
- 1,048,576 input tokens;
- 65,536 output tokens;
- text, image, audio, video, and document input;
- text output;
- dynamic thinking;
- structured outputs;
- context caching;
- search grounding;
- code execution;
- function calling;
- URL context;
- Computer Use in preview.
- Provider
- Family
- Gemini 3 Flash
- Model ID
- gemini-3-flash-preview
- Launch stage
- Public preview · legacy
- Release date
- Dec 17, 2025
- Knowledge cutoff
- Jan 2025
- Output
- Text
02 / Frontier + speed
Pro-Grade Reasoning Without Pro-Tier Latency
Gemini 3 Flash Preview launched around a specific promise: preserve much of Gemini 3’s high-end reasoning while reducing cost and latency enough for interactive and high-frequency applications.
Speed Became Part of Frontier Capability
Before Gemini 3 Flash, developers often treated “frontier reasoning” and “fast production model” as separate categories.
Google explicitly challenged that split.
Gemini 3 Flash was presented as combining Gemini 3’s Pro-grade reasoning foundation with Flash-level latency, efficiency, and cost.
This made it relevant to workloads such as:
- interactive analysis;
- coding assistants;
- rapid agent loops;
- multimodal applications;
- document analysis;
- search-grounded answers;
- UI generation;
- tool-heavy application workflows.
Google also made Gemini 3 Flash the default model in the Gemini app and AI Mode in Search at launch, demonstrating that the model was intended for high-frequency user-facing interactions rather than only offline API work.
The correct evaluation metric is therefore not raw intelligence alone.
Measure:
- accepted-result quality;
- time to first useful answer;
- complete-task latency;
- reasoning-token usage;
- retries;
- tool calls;
- cost per accepted result.
- 01
Reason
Apply Gemini 3 reasoning to complex questions, coding, multimodal analysis, and agentic workflows.
- 02
Respond quickly
Use Flash-class latency for interactive experiences where a strong answer must also arrive quickly.
- 03
Act
Move from reasoning into search, code execution, functions, URLs, and Computer Use.
- 04
Measure the full result
Evaluate quality, latency, tokens, retries, and tools together rather than optimizing a single benchmark.
03 / Agentic coding
The Gemini 3 Flash Preview Was Built for Agentic Coding
Agentic coding was one of the defining launch workloads for Gemini 3 Flash Preview, with Google positioning the model for high-frequency terminal workflows and iterative software-engineering tasks.
Coding Is More Than Generating a Function
Google brought Gemini 3 Flash into Gemini CLI immediately at launch.
The model was designed for workflows where the AI repeatedly:
- reads code;
- searches files;
- reasons about dependencies;
- edits implementations;
- runs commands;
- interprets failures;
- retries;
- validates the final result.
Google reported a 78% SWE-bench Verified score for Gemini 3 Flash in its December 2025 Gemini CLI announcement.
That benchmark number is useful context, but production evaluation should focus on complete engineering trajectories.
Measure:
- first-pass test success;
- number of patch iterations;
- repeated file reads;
- invalid tool calls;
- command failures;
- unwanted edits;
- human intervention;
- latency to accepted implementation.
This page therefore treats Gemini 3 Flash Preview as an agentic-coding baseline, not merely as a generic chat model.
- 01
Inspect
Read repository files, logs, errors, dependencies, and surrounding project context.
- 02
Plan
Choose an implementation path and decide which tools or files are actually needed.
- 03
Execute
Edit code, invoke tools, run commands, and use execution results as evidence.
- 04
Repair
Iterate after failed tests or incorrect assumptions until the task meets acceptance criteria.
04 / Visual reasoning
Code Execution Can Turn Vision Into an Active Investigation
One of Gemini 3 Flash Preview’s most distinctive API features is the ability to combine image reasoning with code execution so the model can zoom, crop, annotate, count, or otherwise manipulate visual input during analysis.
The Model Can Inspect an Image Step by Step
A normal vision model receives an image and produces an interpretation.
Gemini 3 Flash Preview can go further.
Google documents a workflow in which the model:
- forms a plan;
- writes Python;
- executes code against the image;
- zooms or crops relevant regions;
- visually grounds the answer using the manipulated result.
This is useful for:
- dense diagrams;
- screenshots;
- small visual details;
- object counting;
- interface inspection;
- charts;
- spatial reasoning;
- technical images.
Google described Gemini 3 Flash at launch as having its most advanced visual and spatial reasoning for the model class.
The important evaluation question is whether active inspection improves accuracy enough to justify extra execution steps and latency.
- 01
Observe
Interpret the initial image and identify which details need closer inspection.
- 02
Plan
Decide whether cropping, zooming, annotation, counting, or another transformation is necessary.
- 03
Execute
Write and run code that transforms or inspects the visual input.
- 04
Ground
Use the manipulated visual evidence to produce a more precise final answer.
05 / Pricing
Gemini 3 Flash Preview Pricing
Gemini 3 Flash Preview uses low Flash pricing: $0.50 per million text, image, or video input tokens and $3 per million output tokens on the standard paid tier.
Audio Input Has a Separate Rate
Google currently lists Standard pricing at:
- $0.50 per 1M text, image, or video input tokens;
- $1.00 per 1M audio input tokens;
- $3.00 per 1M output tokens, including thinking;
- $0.05 per 1M cached text, image, or video tokens;
- $0.10 per 1M cached audio tokens;
- $1.00 per 1M cached tokens per hour of storage.
Batch and Flex pricing reduce token rates to:
- $0.25 text/image/video input;
- $0.50 audio input;
- $1.50 output.
Priority pricing increases rates to:
- $0.90 text/image/video input;
- $1.80 audio input;
- $5.40 output.
The Gemini Developer API also provides a free tier for Gemini 3 Flash Preview standard token usage.
Thinking tokens are included in output billing, so high-default reasoning can make the effective cost larger than the visible final response suggests.
1M tokens · USD
- Text / image / video input
- $0.50
- Audio input
- $1.00
- Cached text / image / video
- $0.05
- Output + thinking
- $3.00
Example: 20K text input + 4K output
- Input cost
- $0.0100
- Output cost
- $0.0120
- Estimated total before extra thinking
- $0.0220
06 / Multimodal context
1,048,576 Tokens of Multimodal Input
Gemini 3 Flash Preview supports a 1,048,576-token context window and up to 65,536 output tokens across text-centric and multimodal workloads.
The Working Context Can Include Much More Than Text
The model can process:
- text;
- images;
- audio;
- video;
- PDFs and other documents;
- conversation history;
- tool results;
- system instructions;
- retrieved URLs.
Google Cloud documents support for up to 3,000 images per prompt, up to 3,000 document pages per file, long-form video, and audio inputs that can extend to roughly 8.4 hours within token limits.
This makes the model useful for:
- repository analysis;
- document review;
- long video understanding;
- audio summarization;
- multimodal research;
- screenshot-heavy workflows;
- persistent agent sessions.
Large context should still be curated.
Sending irrelevant history or media increases cost and can reduce the model’s ability to focus on the actual task.
Input context
1,048,576
Max output
65,536
Text, images, audio, video, documents, tool results, history, and instructions all share the same working context.
07 / Thinking
High Thinking Is the Default
Gemini 3 Flash Preview supports minimal, low, medium, and high thinking levels, but unlike later stable Flash models it defaults to high.
The Preview Was Tuned Toward Reasoning Depth
Google’s Gemini 3 developer guide documents:
minimal;low;medium;high.
high is the default and uses dynamic thinking.
That means an application that does not explicitly configure reasoning may spend more time and output tokens thinking than a later Flash integration that defaults to medium.
Use:
- minimal for very simple, latency-sensitive requests;
- low for straightforward instruction following and high-throughput work;
- medium for a balanced quality/latency tradeoff;
- high for difficult coding, multimodal reasoning, mathematics, and tool-heavy workflows.
Google retains the older thinking_budget parameter for compatibility but recommends thinking_level. The two cannot be used in the same request.
Interactive workload
Reduce thinking to minimal or low when response latency matters and deeper reasoning does not improve acceptance.
Complex agent workload
Keep medium or high for difficult coding, visual reasoning, and multi-step tool orchestration.
08 / Tools
Built-In Tools and Custom Functions Can Work Together
Gemini 3 introduced richer tool composition, allowing built-in Google tools and developer-defined functions to participate in the same workflow.
One Agent Can Retrieve, Compute, and Act
The Gemini 3 developer documentation lists support for:
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- URL context;
- function calling;
- structured outputs;
- context caching.
Gemini 3 also supports combining built-in tools with custom function calls.
For example, a workflow can use Google Search to identify current information and then call an application-defined function to act on that information.
Structured outputs can be combined with supported tools so the final grounded answer can still conform to a machine-readable schema.
This makes Gemini 3 Flash Preview useful for agents that must move between retrieval, reasoning, computation, and application actions.
- Supported
Google Search
Ground responses in current web information beyond the model’s January 2025 knowledge cutoff.
- Supported
Google Maps
Use Maps grounding in supported Gemini 3 workflows.
- Supported
File Search
Retrieve relevant content from indexed file collections.
- Supported
Code execution
Run code for calculations, verification, and active visual investigation.
- Supported
Function calling
Invoke application-defined tools and combine them with supported built-in tools.
- Supported
Structured outputs
Return schema-constrained machine-readable results.
- Supported
URL context
Read and reason over content from supplied URLs.
- Not listed
Live API
Google Cloud lists Gemini Live API as unsupported for this model.
09 / Computer Use
Computer Use Was Added to Gemini 3 Flash Preview
Google added Computer Use support to Gemini 3 Flash Preview in January 2026, allowing the model to reason over graphical interfaces and act through browser or UI automation workflows.
The Flash Model Can Work Through Interfaces, Not Only APIs
Computer Use is relevant when the system being automated does not expose the required operation through a clean API.
An agent can inspect visual state, decide the next action, interact with controls, observe the result, and continue.
Potential workloads include:
- browser automation;
- application testing;
- repetitive enterprise workflows;
- data entry;
- cross-system operations;
- UI-based research.
The feature remains in preview.
Production systems should evaluate more than successful demos.
Measure incorrect clicks, navigation loops, destructive actions, recovery after unexpected UI state, confirmation requirements, and successful termination.
- 01
Observe
Interpret screenshots and current application state.
- 02
Decide
Reason about the next interaction needed to advance the task.
- 03
Act
Interact with supported interfaces through the Computer Use tool.
- 04
Verify
Inspect the updated state and continue, recover, or stop when the target outcome is reached.
10 / Preview lifecycle
Gemini 3 Flash Preview Is Now a Legacy Preview Model
The endpoint remains available, but Google now labels it as a legacy Flash model and recommends moving production workloads to a newer stable Flash release.
No Shutdown Date Is Announced Yet
Google’s deprecation table currently lists:
- release date: December 17, 2025;
- status: preview model;
- shutdown date: no shutdown date announced;
- recommended replacement:
gemini-3.6-flash.
Google’s general model documentation notes that preview models may be used in production but can have more restrictive rate limits and are normally deprecated with advance notice.
That means the model can still be useful for:
- regression testing;
- compatibility evaluation;
- historical model comparisons;
- applications that have not yet migrated.
It is less attractive as the default choice for a brand-new production system.
Keep the model ID configurable and maintain regression tests so moving to a stable replacement does not require rewriting provider-specific application logic.
- 01
Still available
Google currently lists no shutdown date for gemini-3-flash-preview.
- 02
Legacy classification
Current pricing documentation describes it as a legacy Flash model providing baseline speed and intelligence.
- 03
Stable replacement
Google’s current deprecation table recommends gemini-3.6-flash as the replacement.
- 04
Keep migration cheap
Store model routing and reasoning settings in configuration so a future retirement does not require application-wide code changes.
11 / Migration
Migrating From Gemini 3 Flash Preview
Moving from Gemini 3 Flash Preview to a stable Flash model is not only a model-ID change: reasoning defaults, request configuration, price, and agent behavior should all be re-baselined.
The First Stable Successors Changed the Operating Profile
Google’s Gemini 3.5 Flash migration guide explicitly documents the move from:
gemini-3-flash-preview → gemini-3.5-flash
The guide calls out several changes:
- default thinking moved from
hightomedium; thinking_levelis preferred overthinking_budget;- older sampling overrides such as
temperature,top_p, andtop_kshould be removed from the migration path; - function responses need correct IDs and names;
- prompts should be re-tested for quality, speed, and cost.
Google’s current deprecation table now recommends gemini-3.6-flash as the replacement for the preview endpoint.
That means new migrations should evaluate the current stable Flash family rather than assuming 3.5 is still the best destination.
The economics also differ.
Gemini 3 Flash Preview costs $0.50 input / $3 output for standard text, image, and video usage.
Later stable Flash models can have different prices and reasoning behavior, so migration should compare complete tasks rather than only API compatibility.
- 01
Update the model target
Move away from the legacy preview to an evaluated stable Flash endpoint; Google currently recommends Gemini 3.6 Flash.
- 02
Re-test thinking
The preview defaults to high, while later Flash releases can default to medium or use different supported thinking levels.
- 03
Clean request parameters
Prefer thinking_level, remove stale sampling overrides where migration guidance recommends it, and validate function-response IDs.
- 04
Re-baseline task economics
Compare accepted-result rate, tokens, retries, tool calls, latency, and total cost before switching production traffic.
12 / Evaluation
Gemini 3 Flash Preview Strengths and Limitations
Gemini 3 Flash Preview remains an important Gemini 3 baseline: it introduced frontier-grade Flash reasoning, rich multimodal tooling, visual code execution, and Computer Use, but newer stable Flash generations are now the better default for most production deployments.
Strengths
Frontier reasoning at Flash economics
The model established the Gemini 3 pattern of bringing stronger reasoning and multimodal capability into a faster, lower-cost tier.
High-default reasoning
Dynamic high thinking makes the preview useful for comparing reasoning-heavy behavior against later medium-default Flash models.
Active visual reasoning
Code execution can zoom, crop, annotate, count, and manipulate images as part of step-by-step visual investigation.
Rich agent tool composition
Search, Maps, files, code, URLs, functions, structured outputs, caching, and Computer Use support complex multimodal workflows.
What to consider
Legacy preview status
Google now labels the model as legacy and recommends migrating to a stable Flash model.
High reasoning by default
Unconfigured requests can spend more time and output tokens reasoning than later Flash models that default to medium.
Preview lifecycle risk
No shutdown date is currently announced, but preview endpoints can be deprecated and should not be hard-coded deeply into application architecture.
Newer Flash generations improve production behavior
Gemini 3.5 through 3.8 add stronger coding, efficiency, long-horizon execution, and newer agent capabilities, making them more relevant for new deployments.
Benchmark the original Gemini 3 Flash baseline
Test Gemini 3 Flash Preview against stable Flash models
Replay coding, visual reasoning, multimodal analysis, tool use, browser automation, and long-context prompts to compare quality, latency, thinking cost, tool-call reliability, and migration impact.
Start FreeGemini 3 Flash Preview is now a legacy preview model. Use it as an evaluation or compatibility baseline, and keep production routing ready to move to a stable Gemini Flash endpoint.
Common Questions
What is Gemini 3 Flash Preview?
Gemini 3 Flash Preview is the original Gemini 3 Flash model, launched as a fast and cost-effective model combining frontier-level reasoning, multimodal intelligence, coding, and agentic capabilities.
What is the Gemini 3 Flash Preview model ID?
The model ID is gemini-3-flash-preview.
When was Gemini 3 Flash Preview released?
Google released Gemini 3 Flash Preview on December 17, 2025.
Is Gemini 3 Flash Preview still available?
Yes. Google currently lists the model as a public-preview endpoint with no announced shutdown date, although current pricing documentation describes it as a legacy Flash model.
What model should replace Gemini 3 Flash Preview?
Google’s current Gemini deprecations table recommends gemini-3.6-flash as the replacement. Applications should also evaluate newer stable Flash releases against their own workload requirements.
What is the Gemini 3 Flash Preview context window?
Gemini 3 Flash Preview supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3 Flash Preview support?
The model supports text, images, audio, video, and document input and generates text output.
What is the Gemini 3 Flash Preview knowledge cutoff?
Google documents January 2025 as the knowledge cutoff for Gemini 3 models. Search grounding can be used when current external information is required.
How much does Gemini 3 Flash Preview cost?
Standard paid pricing is $0.50 per 1M text, image, or video input tokens, $1 per 1M audio input tokens, and $3 per 1M output tokens including thinking.
How much does Gemini 3 Flash Preview context caching cost?
Standard cache reads cost $0.05 per 1M text, image, or video tokens and $0.10 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.
Does Gemini 3 Flash Preview have a free tier?
Yes. Google currently lists free standard token usage for Gemini 3 Flash Preview in the Gemini Developer API, subject to provider limits and terms.
What are the Batch prices for Gemini 3 Flash Preview?
Batch pricing is $0.25 per 1M text, image, or video input tokens, $0.50 per 1M audio input tokens, and $1.50 per 1M output tokens.
Does Gemini 3 Flash Preview support thinking?
Yes. It supports minimal, low, medium, and high thinking levels, with high as the default dynamic setting.
Does Gemini 3 Flash Preview support minimal thinking?
Yes. Minimal is supported and is intended to minimize latency for simple chat and high-throughput tasks, although Google notes that the model may still perform very limited reasoning on complex requests.
Can I use thinking_budget with Gemini 3 Flash Preview?
Yes for backward compatibility, but Google recommends thinking_level for more predictable behavior. Do not send thinking_budget and thinking_level together.
Is Gemini 3 Flash Preview good for coding?
Yes. Agentic coding was one of its launch workloads, and Google brought the model to Gemini CLI for high-frequency terminal workflows.
What is code execution with images?
Gemini 3 Flash Preview can reason about an image, write and execute code to crop, zoom, annotate, count, or otherwise manipulate the visual input, then use the transformed evidence to answer more precisely.
Does Gemini 3 Flash Preview support Computer Use?
Yes, in preview. Google added Computer Use support to gemini-3-flash-preview in January 2026.
What tools does Gemini 3 Flash Preview support?
Gemini 3 documentation lists Search grounding, Maps grounding, File Search, code execution, URL context, function calling, structured outputs, and context caching. Computer Use is also supported in preview.
Does Gemini 3 Flash Preview support the Live API?
No. Google Cloud currently lists Gemini Live API as unsupported for this model.
How is Gemini 3 Flash Preview different from Gemini 3.5 Flash?
Gemini 3 Flash Preview is the original legacy preview with high thinking by default and $0.50/$3 standard text pricing. Gemini 3.5 Flash is a stable successor focused more heavily on frontier agentic coding, automatic thought preservation, and production use, with medium thinking by default and higher token pricing.
Should I use Gemini 3 Flash Preview for a new production application?
For a new production application, a stable Gemini Flash model is generally the better default. Gemini 3 Flash Preview remains useful for compatibility, regression testing, historical comparison, or workloads where its specific price and behavior still outperform the migration alternatives.
Model information
Last updated
Specifications, pricing, launch date, preview status, thinking levels, visual reasoning, code execution with images, multimodal function responses, tools, Computer Use, deprecation guidance, and migration behavior on this page are based on official Google Gemini API, Google Cloud, and Google product documentation.