Google Gemini legacy workhorse
Gemini 2.5 Flash
Google’s first hybrid-reasoning Flash model: a mature price-performance workhorse for low-latency reasoning, large-scale multimodal processing, structured extraction, grounded workflows, and tunable thinking.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $0.30
- text/image/video · per 1M
- Cached input
- $0.03
- text/image/video · per 1M
- Output
- $2.50
- per 1M tokens
01 / Overview
What Gemini 2.5 Flash Is
Gemini 2.5 Flash is Google’s first hybrid-reasoning Flash model and one of the defining production models of the Gemini 2.5 generation.
A Mature Price-Performance Workhorse
Google introduced Gemini 2.5 Flash in preview on April 17, 2025 and promoted the stable gemini-2.5-flash endpoint to general availability on June 17, 2025.
The model was designed to combine three properties that previously pulled in different directions:
- low latency;
- strong reasoning;
- production-scale cost efficiency.
Its core innovation was controllable thinking.
Unlike a model that always reasons deeply or never reasons at all, Gemini 2.5 Flash can dynamically decide how much reasoning to use, accept a developer-defined thinking budget, or disable thinking entirely.
That makes it useful for workloads where the same model needs to serve both simple requests and more difficult multi-step tasks.
Google currently describes Gemini 2.5 Flash as a strong price-performance model for:
- large-scale processing;
- low-latency applications;
- high-volume workloads that still require reasoning;
- agentic use cases.
The stable endpoint supports:
- 1,048,576 input tokens;
- 65,536 output tokens;
- text, images, video, and audio input;
- text output;
- thinking;
- caching;
- Search grounding;
- Maps grounding;
- File Search;
- code execution;
- function calling;
- structured outputs;
- URL context;
- Batch, Flex, and Priority inference.
- Provider
- Family
- Gemini 2.5 Flash
- Model ID
- gemini-2.5-flash
- Status
- Stable · legacy access
- GA date
- Jun 17, 2025
- Knowledge cutoff
- Jan 2025
- Output
- Text
02 / Hybrid reasoning
Gemini 2.5 Flash Was Google’s First Fully Hybrid Reasoning Model
The defining feature of Gemini 2.5 Flash is not simply that it can reason—it is that developers can decide how much reasoning is appropriate for each request.
One Model Can Behave Like a Fast Flash or a Thinking Model
Google launched 2.5 Flash as its first fully hybrid reasoning model.
For a simple workload, the application can set the thinking budget to zero and prioritize latency.
For a difficult task, the model can use dynamic thinking or a larger fixed budget.
This was an important design shift.
Instead of routing every easy request to one model and every difficult request to another, developers could keep a single endpoint and control reasoning at request time.
This is useful for mixed workloads such as:
- chat where most questions are simple but some require multi-step analysis;
- extraction pipelines with occasional ambiguous documents;
- routing systems that sometimes need deeper classification logic;
- research workflows where only some requests require search and reasoning;
- agent systems with both deterministic and difficult steps.
For evaluation, compare the same task at several thinking budgets.
Measure:
- accepted-result quality;
- reasoning-token usage;
- visible output length;
- total latency;
- retry rate;
- tool-call accuracy;
- cost per accepted result.
- 01
No thinking
Set thinkingBudget to 0 when low latency matters more than deeper reasoning.
- 02
Fixed budget
Allocate a specific reasoning-token ceiling for workloads where predictable compute is useful.
- 03
Dynamic thinking
Use thinkingBudget = -1 so the model adjusts reasoning depth to request complexity.
- 04
Evaluate the curve
Measure where additional reasoning stops producing enough quality improvement to justify extra latency and cost.
03 / Production workloads
Low-Latency Reasoning for Large-Scale Production
Gemini 2.5 Flash sits between simple high-throughput models and expensive frontier reasoning models, making it useful when a production workload needs both scale and non-trivial reasoning.
Reasoning Without Routing Everything to Pro
Google’s production positioning for 2.5 Flash includes:
- large-scale summarization;
- responsive chat;
- efficient data extraction;
- classification;
- translation;
- intelligent routing;
- agentic workflows.
The model can also process large multimodal inputs, which makes it useful for document-heavy and media-heavy systems rather than only text APIs.
Examples include:
- extracting structured data from long documents;
- analyzing screenshots together with text instructions;
- summarizing long audio or video;
- grounding an answer with Search;
- retrieving from indexed files;
- executing code for calculations;
- calling application-defined functions;
- producing schema-constrained output.
In mature systems, Gemini 2.5 Flash often makes sense as a middle tier.
Very simple tasks can move to Flash-Lite. Difficult long-horizon work can move to newer Gemini 3 Flash or Pro models. The 2.5 Flash tier remains useful where reasoning is necessary but premium-model economics are unnecessary.
- 01
Extract and classify
Process large volumes of documents or records while retaining reasoning for ambiguous cases.
- 02
Respond interactively
Use a Flash-class model for chat and application workflows that still require multi-step thinking.
- 03
Ground and retrieve
Combine Search, Maps, File Search, URLs, and long context with application-specific instructions.
- 04
Escalate selectively
Keep routine reasoning on 2.5 Flash and route only the hardest work to stronger successor models.
04 / Pricing
Gemini 2.5 Flash Pricing
Gemini 2.5 Flash uses a single output price whether reasoning is enabled or disabled: thinking tokens are included in output billing.
Stable Pricing Simplified the Original Preview Model
When Gemini 2.5 Flash first launched, Google used separate thinking and non-thinking prices.
With the stable release, Google simplified the model to one price structure.
Current Standard paid pricing is:
- $0.30 per 1M text, image, or video input tokens;
- $1.00 per 1M audio input tokens;
- $2.50 per 1M output tokens, including thinking;
- $0.03 per 1M cached text, image, or video tokens;
- $0.10 per 1M cached audio tokens;
- $1.00 per 1M cached tokens per hour of storage.
Batch pricing reduces text, image, and video input to $0.15 per 1M tokens and output to $1.25 per 1M tokens.
Priority serving increases standard rates for workloads that need higher-priority capacity.
The Gemini Developer API also lists a free Standard tier, subject to current provider limits.
1M tokens · USD
- Text / image / video input
- $0.30
- Audio input
- $1.00
- Cached text / image / video
- $0.03
- Output + thinking
- $2.50
Example: 20K text input + 4K total output
- Input cost
- $0.0060
- Output cost
- $0.0100
- Estimated total
- $0.0160
05 / Multimodal context
A 1M-Token Multimodal Working Context
Gemini 2.5 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens across text and multimodal production workflows.
Large Context Predates the Gemini 3 Agent Era
A 1M-token window allows the model to receive much more than a user message.
The working context can include:
- long documents;
- images;
- video;
- audio;
- source code;
- conversation history;
- retrieved files;
- URLs;
- system instructions;
- tool results.
That makes Gemini 2.5 Flash useful for document processing, long-context summarization, multimodal extraction, research, and RAG.
The model can accept text, images, video, and audio natively and return text.
Google’s current model specification also lists File Search and URL context, making it possible to combine long native context with retrieval rather than simply sending everything on every request.
For production systems, measure how much of the 1M window is actually useful.
Retrieval, caching, chunking, and selective history can reduce cost and improve signal-to-noise ratio.
Input context
1,048,576
Max output
65,536
The model can combine text, images, video, audio, retrieved data, history, instructions, and tool results inside the same working context.
06 / Thinking budget
Thinking Budget: From Zero to 24,576 Tokens
Gemini 2.5 Flash exposes a direct reasoning-token budget instead of the low/medium/high thinking_level abstraction used by Gemini 3 models.
Fine-Grained Control Is the Core 2.5 Flash Difference
Google documents the following behavior for thinkingBudget:
- range:
0to24576; 0: disable thinking;-1: enable dynamic thinking;- no explicit budget: dynamic thinking by default.
The budget is a ceiling, not necessarily a target. The model can use fewer reasoning tokens when the task does not require the full allowance.
This makes 2.5 Flash useful for empirical reasoning-cost curves.
You can replay the same workload with:
- thinking disabled;
- a small fixed budget;
- a medium budget;
- a large budget;
- dynamic thinking.
Then measure when more reasoning actually improves acceptance.
That is a more granular control surface than later Gemini 3 thinking_level presets.
Latency-sensitive request
Set thinkingBudget to 0 or a small fixed value when deeper reasoning does not materially improve the result.
Complex reasoning request
Use a larger fixed budget or dynamic thinking when multi-step analysis, coding, or tool decisions benefit from more deliberation.
07 / Tools
A Broad Tool Surface for a Pre-Gemini-3 Model
Gemini 2.5 Flash supports a mature production tool set including Search, Maps, File Search, code execution, URLs, functions, structured outputs, and caching.
Ground, Retrieve, Compute, and Act
Google currently lists support for:
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- function calling;
- structured outputs;
- URL context;
- context caching;
- Batch API;
- Flex inference;
- Priority inference.
This gives 2.5 Flash enough capability to operate in grounded application workflows rather than only answer from model memory.
Example patterns include:
- Search → summarize → structured output;
- File Search → extract → validate;
- URL context → compare → classify;
- code execution → calculate → explain;
- function call → observe result → continue.
The base gemini-2.5-flash endpoint does not support the Gemini Live API.
Google provided separate Gemini 2.5 native-audio Live models for real-time voice experiences. Those are different model IDs with different context limits and pricing.
- Supported
Google Search
Ground responses in current web information.
- Supported
Google Maps
Use Maps grounding for location-aware workflows.
- Supported
File Search
Retrieve content from indexed files for RAG and document workflows.
- Supported
Code execution
Run code for calculations, transformations, and verification.
- Supported
Function calling
Invoke application-defined actions and tools.
- Supported
Structured outputs
Return schema-constrained machine-readable results.
- Supported
URL context
Read and reason over supplied URLs.
- Not listed
Live API
The base gemini-2.5-flash model does not support Live API; Google offers separate Gemini 2.5 native-audio Live models.
08 / Tuning
Gemini 2.5 Flash Supports Multiple Tuning Paths on Vertex AI
Gemini 2.5 Flash is not only a prompt-driven API model: Google Cloud supports supervised fine-tuning and preference-based customization for specialized production behavior.
Mature Workloads Can Benefit From Customization
Google Cloud lists Gemini 2.5 Flash among the models that support supervised fine-tuning.
It also lists preference tuning for Gemini 2.5 Flash.
The model-specific Google Cloud documentation additionally lists:
- supervised fine-tuning;
- continuous tuning;
- preference tuning;
- tuning checkpoints.
Potential use cases include:
- domain-specific extraction;
- classification;
- sentiment analysis;
- specialized terminology;
- organization-specific writing behavior;
- domain-specific code generation;
- stable formatting rules.
Tuning is not automatically better than prompting.
Before customizing, establish a prompt-only baseline and measure accuracy, prompt-token consumption, latency, training/maintenance overhead, and regression risk.
For a high-volume model, reducing repeated prompt examples can sometimes be as valuable as improving task accuracy.
- 01
Baseline
Evaluate the untuned model on representative production data and acceptance criteria.
- 02
Tune
Use supervised or preference tuning when stable specialized behavior is difficult to achieve through prompts alone.
- 03
Compare
Measure quality, prompt length, latency, and cost against the original model.
- 04
Monitor
Use checkpoints and regression tests to detect degradation as data and requirements evolve.
09 / Lifecycle
Stable, Not Deprecated—But No Longer the Default for New Projects
Gemini 2.5 Flash remains a stable API model, but Google now limits access to users who have actively used the 2.5 family and recommends newer Gemini models for new projects.
Legacy Workloads and New Projects Have Different Guidance
Google’s Gemini API deprecation page currently states that Gemini 2.5 Flash is not deprecated and will continue to be served until further notice.
However, Google also says access to the 2.5 family is being limited to users who have actively used these models in the past.
For new projects, Google recommends newer models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.
That makes Gemini 2.5 Flash a legacy production model rather than a dead model.
Existing applications may continue using it where migration risk or validated behavior makes that worthwhile.
New applications should generally evaluate current Gemini 3 alternatives first.
There is also a platform-specific lifecycle distinction.
Google Cloud’s Gemini Enterprise Agent Platform documentation states that Gemini 2.5 Flash is being retired there on October 20, 2026. That platform retirement should not be confused with a universal Gemini API shutdown, because the Gemini API deprecation table still lists no shutdown date for the stable endpoint.
- 01
Stable API model
Google’s Gemini API deprecation table does not currently mark gemini-2.5-flash as deprecated.
- 02
Existing-user access
Google is limiting 2.5 access to users who have actively used the family in the past.
- 03
New-project guidance
Google recommends current models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash for new projects.
- 04
Platform-specific retirement
Google Cloud’s Gemini Enterprise Agent Platform lists an October 20, 2026 retirement for Gemini 2.5 Flash on that platform.
10 / Migration to Gemini 3
Migrating Gemini 2.5 Flash Workloads to Gemini 3
The biggest migration difference is conceptual: Gemini 2.5 Flash exposes explicit thinking-token budgets, while Gemini 3 models use reasoning levels and newer agent behavior.
Do More Than Replace the Model ID
A migration should re-evaluate:
- reasoning controls;
- latency;
- output-token consumption;
- tool-call behavior;
- structured-output reliability;
- long-context behavior;
- total task cost.
Gemini 2.5 Flash uses thinkingBudget.
Gemini 3 uses thinking_level for the primary reasoning control surface.
That means an old setting such as:
thinkingBudget: 0;thinkingBudget: 1024;thinkingBudget: -1;
does not translate one-to-one into low, medium, or high.
Build a regression set and map each workload empirically.
Google currently recommends newer Gemini 3 models for new projects.
A useful routing comparison is:
- Gemini 3.5 Flash-Lite for cheaper, high-throughput routine work;
- Gemini 3.8 Flash for stronger coding, agents, and difficult multi-step workloads.
If an existing 2.5 Flash workflow already performs reliably, migration should still be measured rather than assumed.
- 01
Map thinking controls
Replace explicit 2.5 thinking budgets with evaluated Gemini 3 thinking levels rather than guessing an equivalent.
- 02
Replay real prompts
Use production-like extraction, multimodal, RAG, coding, and tool-use cases to compare behavior.
- 03
Measure agent trajectories
Count retries, malformed calls, repeated actions, and total steps because newer Gemini models can change orchestration behavior.
- 04
Compare cost per accepted result
Include thinking tokens, latency, retries, fallbacks, and tool overhead rather than comparing published token prices alone.
11 / Evaluation
Gemini 2.5 Flash Strengths and Limitations
Gemini 2.5 Flash remains a capable production baseline with unusually fine-grained reasoning control, but it now belongs to the legacy side of Google’s model portfolio.
Strengths
Fine-grained reasoning control
thinkingBudget can be disabled, fixed from 0 to 24,576 tokens, or set to dynamic reasoning.
Strong price-performance
$0.30 text/image/video input and $2.50 output pricing supports reasoning-heavy workloads without Pro-tier token costs.
Mature multimodal tool stack
A 1M context, Search, Maps, File Search, code execution, functions, structured outputs, URLs, and caching support broad production use.
Customization support
Vertex AI supports supervised fine-tuning, continuous tuning, preference tuning, and checkpoints for specialized workloads.
What to consider
Legacy access
Google now limits 2.5 access to existing active users and recommends Gemini 3 models for new projects.
Older reasoning API
thinkingBudget is more granular but differs from Gemini 3 thinking_level, adding migration work.
No Live API on the base model
Real-time native-audio experiences use separate Gemini 2.5 Live model IDs rather than gemini-2.5-flash.
Newer models improve agentic execution
Gemini 3 generations provide stronger modern coding, computer-use, and long-horizon agent behavior for many new workloads.
Benchmark the mature reasoning baseline
Test Gemini 2.5 Flash against newer Gemini models
Replay reasoning, extraction, multimodal processing, grounded search, RAG, tool-use, and high-volume production tasks to compare accepted-result quality, thinking tokens, latency, fallback rate, and total cost.
Start FreeGemini 2.5 Flash remains useful for established workloads, but new projects should benchmark current Gemini 3 models because access to 2.5 is increasingly oriented toward existing users.
Common Questions
What is Gemini 2.5 Flash?
Gemini 2.5 Flash is Google’s first hybrid-reasoning Flash model, designed for strong price-performance across low-latency, high-volume, multimodal, reasoning, and agentic workloads.
What is the Gemini 2.5 Flash model ID?
The stable model ID is gemini-2.5-flash.
When was Gemini 2.5 Flash released?
Google introduced Gemini 2.5 Flash in preview on April 17, 2025 and released the stable generally available model on June 17, 2025.
Is Gemini 2.5 Flash deprecated?
No. Google’s current Gemini API deprecation page says the stable 2.5 models are not deprecated and will continue to be served until further notice, but access is increasingly limited to existing active users.
Can new users access Gemini 2.5 Flash?
Google says access to Gemini 2.5 models is being limited to users who have actively used them in the past. For new projects, Google recommends newer Gemini 3 models.
What is the Gemini 2.5 Flash context window?
Gemini 2.5 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 2.5 Flash support?
The model accepts text, images, video, and audio and generates text output.
What is the Gemini 2.5 Flash knowledge cutoff?
Google documents January 2025 as the knowledge cutoff.
How much does Gemini 2.5 Flash cost?
Standard paid pricing is $0.30 per 1M text, image, or video input tokens, $1 per 1M audio input tokens, and $2.50 per 1M output tokens including thinking.
How much does Gemini 2.5 Flash context caching cost?
Standard cache reads cost $0.03 per 1M text, image, or video tokens and $0.10 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.
What are Gemini 2.5 Flash Batch prices?
Google lists Batch input pricing at $0.15 per 1M text, image, or video tokens, with discounted output pricing relative to Standard serving.
Does Gemini 2.5 Flash have a free tier?
Yes. Google currently lists free Standard token usage in the Gemini Developer API, subject to provider limits and terms.
Does Gemini 2.5 Flash support thinking?
Yes. It was Google’s first fully hybrid reasoning model in the Flash family.
What is the Gemini 2.5 Flash thinking budget?
The supported thinkingBudget range is 0 to 24,576 tokens. Setting it to 0 disables thinking, while -1 enables dynamic thinking and is the default behavior when no fixed budget is supplied.
Can Gemini 2.5 Flash run with thinking disabled?
Yes. Set thinkingBudget to 0 to disable reasoning and prioritize lower latency and output-token usage.
What does dynamic thinking mean on Gemini 2.5 Flash?
With thinkingBudget set to -1, the model dynamically chooses how much reasoning to use based on request complexity rather than consuming a fixed budget.
Does Gemini 2.5 Flash support File Search?
Yes. Google’s current model specification lists File Search as supported.
What tools does Gemini 2.5 Flash support?
Google lists Search grounding, Maps grounding, File Search, code execution, function calling, structured outputs, URL context, and context caching.
Does Gemini 2.5 Flash support the Live API?
No. The base gemini-2.5-flash endpoint does not support Live API. Google provides separate Gemini 2.5 native-audio Live models for real-time audio experiences.
Does Gemini 2.5 Flash support tuning?
Yes on Vertex AI. Google Cloud lists supervised fine-tuning, continuous tuning, preference tuning, and tuning checkpoints for Gemini 2.5 Flash.
Is Gemini 2.5 Flash being retired on October 20, 2026?
Google Cloud’s Gemini Enterprise Agent Platform lists October 20, 2026 as a retirement date on that platform. The general Gemini API deprecation page currently lists no shutdown date for the stable gemini-2.5-flash endpoint, so the platform-specific retirement should not be treated as a universal API shutdown.
How is Gemini 2.5 Flash different from Gemini 3 Flash models?
Gemini 2.5 Flash uses explicit thinkingBudget control from 0 to 24,576 tokens, while Gemini 3 models primarily use thinking levels. Newer Gemini 3 Flash models also focus more strongly on modern agentic coding, computer use, and long-horizon execution.
Should I use Gemini 2.5 Flash for a new application?
For a new application, Google recommends newer models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash. Gemini 2.5 Flash remains most relevant for established workloads that already depend on its validated behavior, pricing, tuning, or fine-grained thinking-budget control.
Model information
Last updated
Specifications, pricing, context limits, thinking budgets, multimodal input, tools, tuning, release status, access restrictions, lifecycle notes, and migration guidance on this page are based on official Google Gemini API, Google Cloud, and Google developer documentation.