Google Gemini legacy workhorse

Gemini 2.5 Flash

Google’s first hybrid-reasoning Flash model: a mature price-performance workhorse for low-latency reasoning, large-scale multimodal processing, structured extraction, grounded workflows, and tunable thinking.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.30
text/image/video · per 1M
Cached input
$0.03
text/image/video · per 1M
Output
$2.50
per 1M tokens

01 / Overview

What Gemini 2.5 Flash Is

Gemini 2.5 Flash is Google’s first hybrid-reasoning Flash model and one of the defining production models of the Gemini 2.5 generation.

A Mature Price-Performance Workhorse

Google introduced Gemini 2.5 Flash in preview on April 17, 2025 and promoted the stable gemini-2.5-flash endpoint to general availability on June 17, 2025.

The model was designed to combine three properties that previously pulled in different directions:

  • low latency;
  • strong reasoning;
  • production-scale cost efficiency.

Its core innovation was controllable thinking.

Unlike a model that always reasons deeply or never reasons at all, Gemini 2.5 Flash can dynamically decide how much reasoning to use, accept a developer-defined thinking budget, or disable thinking entirely.

That makes it useful for workloads where the same model needs to serve both simple requests and more difficult multi-step tasks.

Google currently describes Gemini 2.5 Flash as a strong price-performance model for:

  • large-scale processing;
  • low-latency applications;
  • high-volume workloads that still require reasoning;
  • agentic use cases.

The stable endpoint supports:

  • 1,048,576 input tokens;
  • 65,536 output tokens;
  • text, images, video, and audio input;
  • text output;
  • thinking;
  • caching;
  • Search grounding;
  • Maps grounding;
  • File Search;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • Batch, Flex, and Priority inference.
Model profile
Provider
Google
Family
Gemini 2.5 Flash
Model ID
gemini-2.5-flash
Status
Stable · legacy access
GA date
Jun 17, 2025
Knowledge cutoff
Jan 2025
Output
Text

02 / Hybrid reasoning

Gemini 2.5 Flash Was Google’s First Fully Hybrid Reasoning Model

The defining feature of Gemini 2.5 Flash is not simply that it can reason—it is that developers can decide how much reasoning is appropriate for each request.

One Model Can Behave Like a Fast Flash or a Thinking Model

Google launched 2.5 Flash as its first fully hybrid reasoning model.

For a simple workload, the application can set the thinking budget to zero and prioritize latency.

For a difficult task, the model can use dynamic thinking or a larger fixed budget.

This was an important design shift.

Instead of routing every easy request to one model and every difficult request to another, developers could keep a single endpoint and control reasoning at request time.

This is useful for mixed workloads such as:

  • chat where most questions are simple but some require multi-step analysis;
  • extraction pipelines with occasional ambiguous documents;
  • routing systems that sometimes need deeper classification logic;
  • research workflows where only some requests require search and reasoning;
  • agent systems with both deterministic and difficult steps.

For evaluation, compare the same task at several thinking budgets.

Measure:

  • accepted-result quality;
  • reasoning-token usage;
  • visible output length;
  • total latency;
  • retry rate;
  • tool-call accuracy;
  • cost per accepted result.
Reasoning can be switched and budgeted
  1. 01

    No thinking

    Set thinkingBudget to 0 when low latency matters more than deeper reasoning.

  2. 02

    Fixed budget

    Allocate a specific reasoning-token ceiling for workloads where predictable compute is useful.

  3. 03

    Dynamic thinking

    Use thinkingBudget = -1 so the model adjusts reasoning depth to request complexity.

  4. 04

    Evaluate the curve

    Measure where additional reasoning stops producing enough quality improvement to justify extra latency and cost.

03 / Production workloads

Low-Latency Reasoning for Large-Scale Production

Gemini 2.5 Flash sits between simple high-throughput models and expensive frontier reasoning models, making it useful when a production workload needs both scale and non-trivial reasoning.

Reasoning Without Routing Everything to Pro

Google’s production positioning for 2.5 Flash includes:

  • large-scale summarization;
  • responsive chat;
  • efficient data extraction;
  • classification;
  • translation;
  • intelligent routing;
  • agentic workflows.

The model can also process large multimodal inputs, which makes it useful for document-heavy and media-heavy systems rather than only text APIs.

Examples include:

  • extracting structured data from long documents;
  • analyzing screenshots together with text instructions;
  • summarizing long audio or video;
  • grounding an answer with Search;
  • retrieving from indexed files;
  • executing code for calculations;
  • calling application-defined functions;
  • producing schema-constrained output.

In mature systems, Gemini 2.5 Flash often makes sense as a middle tier.

Very simple tasks can move to Flash-Lite. Difficult long-horizon work can move to newer Gemini 3 Flash or Pro models. The 2.5 Flash tier remains useful where reasoning is necessary but premium-model economics are unnecessary.

Price-performance workloads
  1. 01

    Extract and classify

    Process large volumes of documents or records while retaining reasoning for ambiguous cases.

  2. 02

    Respond interactively

    Use a Flash-class model for chat and application workflows that still require multi-step thinking.

  3. 03

    Ground and retrieve

    Combine Search, Maps, File Search, URLs, and long context with application-specific instructions.

  4. 04

    Escalate selectively

    Keep routine reasoning on 2.5 Flash and route only the hardest work to stronger successor models.

04 / Pricing

Gemini 2.5 Flash Pricing

Gemini 2.5 Flash uses a single output price whether reasoning is enabled or disabled: thinking tokens are included in output billing.

Stable Pricing Simplified the Original Preview Model

When Gemini 2.5 Flash first launched, Google used separate thinking and non-thinking prices.

With the stable release, Google simplified the model to one price structure.

Current Standard paid pricing is:

  • $0.30 per 1M text, image, or video input tokens;
  • $1.00 per 1M audio input tokens;
  • $2.50 per 1M output tokens, including thinking;
  • $0.03 per 1M cached text, image, or video tokens;
  • $0.10 per 1M cached audio tokens;
  • $1.00 per 1M cached tokens per hour of storage.

Batch pricing reduces text, image, and video input to $0.15 per 1M tokens and output to $1.25 per 1M tokens.

Priority serving increases standard rates for workloads that need higher-priority capacity.

The Gemini Developer API also lists a free Standard tier, subject to current provider limits.

Standard and discounted inference

1M tokens · USD

Text / image / video input
$0.30
Audio input
$1.00
Cached text / image / video
$0.03
Output + thinking
$2.50

Example: 20K text input + 4K total output

Input cost
$0.0060
Output cost
$0.0100
Estimated total
$0.0160

05 / Multimodal context

A 1M-Token Multimodal Working Context

Gemini 2.5 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens across text and multimodal production workflows.

Large Context Predates the Gemini 3 Agent Era

A 1M-token window allows the model to receive much more than a user message.

The working context can include:

  • long documents;
  • images;
  • video;
  • audio;
  • source code;
  • conversation history;
  • retrieved files;
  • URLs;
  • system instructions;
  • tool results.

That makes Gemini 2.5 Flash useful for document processing, long-context summarization, multimodal extraction, research, and RAG.

The model can accept text, images, video, and audio natively and return text.

Google’s current model specification also lists File Search and URL context, making it possible to combine long native context with retrieval rather than simply sending everything on every request.

For production systems, measure how much of the 1M window is actually useful.

Retrieval, caching, chunking, and selective history can reduce cost and improve signal-to-noise ratio.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

The model can combine text, images, video, audio, retrieved data, history, instructions, and tool results inside the same working context.

06 / Thinking budget

Thinking Budget: From Zero to 24,576 Tokens

Gemini 2.5 Flash exposes a direct reasoning-token budget instead of the low/medium/high thinking_level abstraction used by Gemini 3 models.

Fine-Grained Control Is the Core 2.5 Flash Difference

Google documents the following behavior for thinkingBudget:

  • range: 0 to 24576;
  • 0: disable thinking;
  • -1: enable dynamic thinking;
  • no explicit budget: dynamic thinking by default.

The budget is a ceiling, not necessarily a target. The model can use fewer reasoning tokens when the task does not require the full allowance.

This makes 2.5 Flash useful for empirical reasoning-cost curves.

You can replay the same workload with:

  • thinking disabled;
  • a small fixed budget;
  • a medium budget;
  • a large budget;
  • dynamic thinking.

Then measure when more reasoning actually improves acceptance.

That is a more granular control surface than later Gemini 3 thinking_level presets.

0–24,576 reasoning tokens
0 · off1K · light8K · deeper24.6K · maxdynamic · default

Latency-sensitive request

Set thinkingBudget to 0 or a small fixed value when deeper reasoning does not materially improve the result.

Complex reasoning request

Use a larger fixed budget or dynamic thinking when multi-step analysis, coding, or tool decisions benefit from more deliberation.

07 / Tools

A Broad Tool Surface for a Pre-Gemini-3 Model

Gemini 2.5 Flash supports a mature production tool set including Search, Maps, File Search, code execution, URLs, functions, structured outputs, and caching.

Ground, Retrieve, Compute, and Act

Google currently lists support for:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • context caching;
  • Batch API;
  • Flex inference;
  • Priority inference.

This gives 2.5 Flash enough capability to operate in grounded application workflows rather than only answer from model memory.

Example patterns include:

  • Search → summarize → structured output;
  • File Search → extract → validate;
  • URL context → compare → classify;
  • code execution → calculate → explain;
  • function call → observe result → continue.

The base gemini-2.5-flash endpoint does not support the Gemini Live API.

Google provided separate Gemini 2.5 native-audio Live models for real-time voice experiences. Those are different model IDs with different context limits and pricing.

Supported capabilities
  • Google Search

    Ground responses in current web information.

    Supported
  • Google Maps

    Use Maps grounding for location-aware workflows.

    Supported
  • File Search

    Retrieve content from indexed files for RAG and document workflows.

    Supported
  • Code execution

    Run code for calculations, transformations, and verification.

    Supported
  • Function calling

    Invoke application-defined actions and tools.

    Supported
  • Structured outputs

    Return schema-constrained machine-readable results.

    Supported
  • URL context

    Read and reason over supplied URLs.

    Supported
  • Live API

    The base gemini-2.5-flash model does not support Live API; Google offers separate Gemini 2.5 native-audio Live models.

    Not listed

08 / Tuning

Gemini 2.5 Flash Supports Multiple Tuning Paths on Vertex AI

Gemini 2.5 Flash is not only a prompt-driven API model: Google Cloud supports supervised fine-tuning and preference-based customization for specialized production behavior.

Mature Workloads Can Benefit From Customization

Google Cloud lists Gemini 2.5 Flash among the models that support supervised fine-tuning.

It also lists preference tuning for Gemini 2.5 Flash.

The model-specific Google Cloud documentation additionally lists:

  • supervised fine-tuning;
  • continuous tuning;
  • preference tuning;
  • tuning checkpoints.

Potential use cases include:

  • domain-specific extraction;
  • classification;
  • sentiment analysis;
  • specialized terminology;
  • organization-specific writing behavior;
  • domain-specific code generation;
  • stable formatting rules.

Tuning is not automatically better than prompting.

Before customizing, establish a prompt-only baseline and measure accuracy, prompt-token consumption, latency, training/maintenance overhead, and regression risk.

For a high-volume model, reducing repeated prompt examples can sometimes be as valuable as improving task accuracy.

Customization options
  1. 01

    Baseline

    Evaluate the untuned model on representative production data and acceptance criteria.

  2. 02

    Tune

    Use supervised or preference tuning when stable specialized behavior is difficult to achieve through prompts alone.

  3. 03

    Compare

    Measure quality, prompt length, latency, and cost against the original model.

  4. 04

    Monitor

    Use checkpoints and regression tests to detect degradation as data and requirements evolve.

09 / Lifecycle

Stable, Not Deprecated—But No Longer the Default for New Projects

Gemini 2.5 Flash remains a stable API model, but Google now limits access to users who have actively used the 2.5 family and recommends newer Gemini models for new projects.

Legacy Workloads and New Projects Have Different Guidance

Google’s Gemini API deprecation page currently states that Gemini 2.5 Flash is not deprecated and will continue to be served until further notice.

However, Google also says access to the 2.5 family is being limited to users who have actively used these models in the past.

For new projects, Google recommends newer models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash.

That makes Gemini 2.5 Flash a legacy production model rather than a dead model.

Existing applications may continue using it where migration risk or validated behavior makes that worthwhile.

New applications should generally evaluate current Gemini 3 alternatives first.

There is also a platform-specific lifecycle distinction.

Google Cloud’s Gemini Enterprise Agent Platform documentation states that Gemini 2.5 Flash is being retired there on October 20, 2026. That platform retirement should not be confused with a universal Gemini API shutdown, because the Gemini API deprecation table still lists no shutdown date for the stable endpoint.

Legacy-access status
  1. 01

    Stable API model

    Google’s Gemini API deprecation table does not currently mark gemini-2.5-flash as deprecated.

  2. 02

    Existing-user access

    Google is limiting 2.5 access to users who have actively used the family in the past.

  3. 03

    New-project guidance

    Google recommends current models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash for new projects.

  4. 04

    Platform-specific retirement

    Google Cloud’s Gemini Enterprise Agent Platform lists an October 20, 2026 retirement for Gemini 2.5 Flash on that platform.

10 / Migration to Gemini 3

Migrating Gemini 2.5 Flash Workloads to Gemini 3

The biggest migration difference is conceptual: Gemini 2.5 Flash exposes explicit thinking-token budgets, while Gemini 3 models use reasoning levels and newer agent behavior.

Do More Than Replace the Model ID

A migration should re-evaluate:

  • reasoning controls;
  • latency;
  • output-token consumption;
  • tool-call behavior;
  • structured-output reliability;
  • long-context behavior;
  • total task cost.

Gemini 2.5 Flash uses thinkingBudget.

Gemini 3 uses thinking_level for the primary reasoning control surface.

That means an old setting such as:

  • thinkingBudget: 0;
  • thinkingBudget: 1024;
  • thinkingBudget: -1;

does not translate one-to-one into low, medium, or high.

Build a regression set and map each workload empirically.

Google currently recommends newer Gemini 3 models for new projects.

A useful routing comparison is:

  • Gemini 3.5 Flash-Lite for cheaper, high-throughput routine work;
  • Gemini 3.8 Flash for stronger coding, agents, and difficult multi-step workloads.

If an existing 2.5 Flash workflow already performs reliably, migration should still be measured rather than assumed.

Migration checklist
  1. 01

    Map thinking controls

    Replace explicit 2.5 thinking budgets with evaluated Gemini 3 thinking levels rather than guessing an equivalent.

  2. 02

    Replay real prompts

    Use production-like extraction, multimodal, RAG, coding, and tool-use cases to compare behavior.

  3. 03

    Measure agent trajectories

    Count retries, malformed calls, repeated actions, and total steps because newer Gemini models can change orchestration behavior.

  4. 04

    Compare cost per accepted result

    Include thinking tokens, latency, retries, fallbacks, and tool overhead rather than comparing published token prices alone.

11 / Evaluation

Gemini 2.5 Flash Strengths and Limitations

Gemini 2.5 Flash remains a capable production baseline with unusually fine-grained reasoning control, but it now belongs to the legacy side of Google’s model portfolio.

Strengths

  • Fine-grained reasoning control

    thinkingBudget can be disabled, fixed from 0 to 24,576 tokens, or set to dynamic reasoning.

  • Strong price-performance

    $0.30 text/image/video input and $2.50 output pricing supports reasoning-heavy workloads without Pro-tier token costs.

  • Mature multimodal tool stack

    A 1M context, Search, Maps, File Search, code execution, functions, structured outputs, URLs, and caching support broad production use.

  • Customization support

    Vertex AI supports supervised fine-tuning, continuous tuning, preference tuning, and checkpoints for specialized workloads.

What to consider

  • Legacy access

    Google now limits 2.5 access to existing active users and recommends Gemini 3 models for new projects.

  • Older reasoning API

    thinkingBudget is more granular but differs from Gemini 3 thinking_level, adding migration work.

  • No Live API on the base model

    Real-time native-audio experiences use separate Gemini 2.5 Live model IDs rather than gemini-2.5-flash.

  • Newer models improve agentic execution

    Gemini 3 generations provide stronger modern coding, computer-use, and long-horizon agent behavior for many new workloads.

Benchmark the mature reasoning baseline

Test Gemini 2.5 Flash against newer Gemini models

Replay reasoning, extraction, multimodal processing, grounded search, RAG, tool-use, and high-volume production tasks to compare accepted-result quality, thinking tokens, latency, fallback rate, and total cost.

Start Free

Gemini 2.5 Flash remains useful for established workloads, but new projects should benchmark current Gemini 3 models because access to 2.5 is increasingly oriented toward existing users.

Common Questions

What is Gemini 2.5 Flash?

Gemini 2.5 Flash is Google’s first hybrid-reasoning Flash model, designed for strong price-performance across low-latency, high-volume, multimodal, reasoning, and agentic workloads.

What is the Gemini 2.5 Flash model ID?

The stable model ID is gemini-2.5-flash.

When was Gemini 2.5 Flash released?

Google introduced Gemini 2.5 Flash in preview on April 17, 2025 and released the stable generally available model on June 17, 2025.

Is Gemini 2.5 Flash deprecated?

No. Google’s current Gemini API deprecation page says the stable 2.5 models are not deprecated and will continue to be served until further notice, but access is increasingly limited to existing active users.

Can new users access Gemini 2.5 Flash?

Google says access to Gemini 2.5 models is being limited to users who have actively used them in the past. For new projects, Google recommends newer Gemini 3 models.

What is the Gemini 2.5 Flash context window?

Gemini 2.5 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 2.5 Flash support?

The model accepts text, images, video, and audio and generates text output.

What is the Gemini 2.5 Flash knowledge cutoff?

Google documents January 2025 as the knowledge cutoff.

How much does Gemini 2.5 Flash cost?

Standard paid pricing is $0.30 per 1M text, image, or video input tokens, $1 per 1M audio input tokens, and $2.50 per 1M output tokens including thinking.

How much does Gemini 2.5 Flash context caching cost?

Standard cache reads cost $0.03 per 1M text, image, or video tokens and $0.10 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.

What are Gemini 2.5 Flash Batch prices?

Google lists Batch input pricing at $0.15 per 1M text, image, or video tokens, with discounted output pricing relative to Standard serving.

Does Gemini 2.5 Flash have a free tier?

Yes. Google currently lists free Standard token usage in the Gemini Developer API, subject to provider limits and terms.

Does Gemini 2.5 Flash support thinking?

Yes. It was Google’s first fully hybrid reasoning model in the Flash family.

What is the Gemini 2.5 Flash thinking budget?

The supported thinkingBudget range is 0 to 24,576 tokens. Setting it to 0 disables thinking, while -1 enables dynamic thinking and is the default behavior when no fixed budget is supplied.

Can Gemini 2.5 Flash run with thinking disabled?

Yes. Set thinkingBudget to 0 to disable reasoning and prioritize lower latency and output-token usage.

What does dynamic thinking mean on Gemini 2.5 Flash?

With thinkingBudget set to -1, the model dynamically chooses how much reasoning to use based on request complexity rather than consuming a fixed budget.

Does Gemini 2.5 Flash support File Search?

Yes. Google’s current model specification lists File Search as supported.

What tools does Gemini 2.5 Flash support?

Google lists Search grounding, Maps grounding, File Search, code execution, function calling, structured outputs, URL context, and context caching.

Does Gemini 2.5 Flash support the Live API?

No. The base gemini-2.5-flash endpoint does not support Live API. Google provides separate Gemini 2.5 native-audio Live models for real-time audio experiences.

Does Gemini 2.5 Flash support tuning?

Yes on Vertex AI. Google Cloud lists supervised fine-tuning, continuous tuning, preference tuning, and tuning checkpoints for Gemini 2.5 Flash.

Is Gemini 2.5 Flash being retired on October 20, 2026?

Google Cloud’s Gemini Enterprise Agent Platform lists October 20, 2026 as a retirement date on that platform. The general Gemini API deprecation page currently lists no shutdown date for the stable gemini-2.5-flash endpoint, so the platform-specific retirement should not be treated as a universal API shutdown.

How is Gemini 2.5 Flash different from Gemini 3 Flash models?

Gemini 2.5 Flash uses explicit thinkingBudget control from 0 to 24,576 tokens, while Gemini 3 models primarily use thinking levels. Newer Gemini 3 Flash models also focus more strongly on modern agentic coding, computer use, and long-horizon execution.

Should I use Gemini 2.5 Flash for a new application?

For a new application, Google recommends newer models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash. Gemini 2.5 Flash remains most relevant for established workloads that already depend on its validated behavior, pricing, tuning, or fine-grained thinking-budget control.

Model information

Last updated

Specifications, pricing, context limits, thinking budgets, multimodal input, tools, tuning, release status, access restrictions, lifecycle notes, and migration guidance on this page are based on official Google Gemini API, Google Cloud, and Google developer documentation.

Gemini 2.5 Flash — Pricing, 1M Context, Thinking Budget & Tools | EidoStack