Google Gemini legacy efficiency model

Gemini 2.5 Flash-Lite

Google’s smallest and most cost-efficient Gemini 2.5 model for classification, extraction, translation, routing, multimodal preprocessing, and other high-volume workloads where speed and budget matter more than deep reasoning.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.10
text/image/video · per 1M
Cached input
$0.01
text/image/video · per 1M
Output
$0.40
per 1M tokens

01 / Overview

What Gemini 2.5 Flash-Lite Is

Gemini 2.5 Flash-Lite is the smallest, fastest, and most cost-efficient model in Google’s Gemini 2.5 family, designed for lightweight tasks that need to run at very high volume.

The Efficiency Floor of the Gemini 2.5 Family

Google introduced Gemini 2.5 Flash-Lite in preview on June 17, 2025 and released the stable gemini-2.5-flash-lite endpoint as generally available on July 22, 2025.

At launch, Google described it as its fastest and most cost-efficient Gemini 2.5 model.

The intended workload is different from Gemini 2.5 Flash.

Flash is the reasoning-oriented price-performance workhorse. Flash-Lite is optimized for the requests where deep reasoning is unnecessary most of the time.

Google specifically highlights:

  • classification;
  • simple data extraction;
  • translation;
  • intelligent routing;
  • high-frequency lightweight tasks;
  • latency-sensitive applications.

Despite the Lite name, the model remains multimodal and supports a large context window.

The stable endpoint accepts text, images, video, audio, and PDFs and generates text.

It supports:

  • 1,048,576 input tokens;
  • 65,536 output tokens;
  • optional thinking;
  • caching;
  • Search grounding;
  • Maps grounding;
  • File Search;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • Batch, Flex, and Priority inference.
Model profile
Provider
Google
Family
Gemini 2.5 Flash-Lite
Model ID
gemini-2.5-flash-lite
Status
Stable · legacy access
GA date
Jul 22, 2025
Knowledge cutoff
Jan 2025
Output
Text

02 / Cost & speed

Built for the Cheapest Acceptable Result at Scale

Gemini 2.5 Flash-Lite is optimized around a practical production constraint: most AI traffic does not need the strongest available model.

Scale Changes What “Best Model” Means

For a single request, the difference between a cheap and expensive model may appear small.

At millions of requests, every input token, output token, retry, and extra second becomes infrastructure cost.

Google reported that Gemini 2.5 Flash-Lite was approximately 1.5× faster than Gemini 2.0 Flash at launch while also reducing cost.

That makes it a natural candidate for:

  • background processing;
  • bulk classification;
  • metadata extraction;
  • translation queues;
  • routing decisions;
  • preprocessing before a larger model;
  • narrow subagents;
  • validation and enrichment jobs.

The correct metric is not “can the model answer the task?”

It is:

Can the model answer the task correctly often enough that the total cost remains lower after retries and fallbacks?

Track:

  • accepted-result rate;
  • first-pass success;
  • p50 and p95 latency;
  • output-token usage;
  • retry rate;
  • fallback rate;
  • cost per completed task.
Efficiency-first design
  1. 01

    Start cheap

    Route predictable, measurable workloads to the lowest-cost model that can plausibly pass.

  2. 02

    Validate

    Use schemas, business rules, confidence checks, or deterministic tests to evaluate the result.

  3. 03

    Escalate

    Send only failed or difficult cases to Gemini 2.5 Flash, Gemini 3, or another stronger model.

  4. 04

    Measure blended cost

    Include retries, validation, fallbacks, and human review when calculating real savings.

03 / High-volume workloads

Classification, Extraction, Translation, and Routing

Gemini 2.5 Flash-Lite is most compelling when the task is narrow, repeatable, and easy to validate at machine scale.

Lightweight Does Not Mean Trivial

A model can be small and still handle valuable production work.

Typical Flash-Lite workloads include:

  • classify support tickets;
  • detect intent;
  • extract names, dates, amounts, and entities;
  • translate product or help-center content;
  • normalize user input;
  • summarize short or medium documents;
  • assign records to queues;
  • route a request to the correct model or tool;
  • generate structured JSON;
  • enrich search indexes;
  • tag images or documents;
  • preprocess long content before a stronger model.

Multimodal input broadens these use cases beyond text.

For example, an extraction pipeline can process a PDF, a scanned image, or an audio transcript without changing to a separate model family.

Use workload-specific metrics.

For extraction, measure field precision and recall.

For classification, measure confusion between categories.

For translation, measure terminology consistency and preservation of numbers and names.

For routing, measure downstream task success rather than only route accuracy.

Routine production tasks
  1. 01

    Classify

    Assign high-volume inputs to categories, queues, policies, intents, or downstream workflows.

  2. 02

    Extract

    Pull entities, values, metadata, and structured fields from text, images, PDFs, or transcripts.

  3. 03

    Translate

    Process repetitive multilingual content where unit cost and throughput matter.

  4. 04

    Route

    Choose the next tool, workflow, or stronger model before expensive reasoning begins.

04 / Pricing

Gemini 2.5 Flash-Lite Pricing

Gemini 2.5 Flash-Lite is one of the least expensive Gemini API models: standard text, image, and video input costs $0.10 per million tokens and output costs $0.40 per million tokens.

Batch and Flex Cut the Core Token Rates in Half

Current Standard paid pricing is:

  • $0.10 per 1M text, image, or video input tokens;
  • $0.30 per 1M audio input tokens;
  • $0.40 per 1M output tokens, including thinking;
  • $0.01 per 1M cached text, image, or video tokens;
  • $0.03 per 1M cached audio tokens;
  • $1.00 per 1M cached tokens per hour of storage.

Batch pricing is:

  • $0.05 per 1M text, image, or video input tokens;
  • $0.15 per 1M audio input tokens;
  • $0.20 per 1M output tokens.

Flex uses the same $0.05 / $0.20 text pricing as Batch.

Priority inference costs more:

  • $0.18 per 1M text, image, or video input tokens;
  • $0.54 per 1M audio input tokens;
  • $0.72 per 1M output tokens.

This structure makes Flash-Lite especially attractive for offline pipelines that can use Batch or Flex.

Ultra-low token pricing

1M tokens · USD

Text / image / video input
$0.10
Audio input
$0.30
Cached text / image / video
$0.01
Output + thinking
$0.40

Example: 20K text input + 4K output

Input cost
$0.0020
Output cost
$0.0016
Estimated total
$0.0036

05 / Multimodal context

A 1M-Token Context Window at Lite Pricing

Gemini 2.5 Flash-Lite keeps the same 1,048,576-token input limit and 65,536-token maximum output used by larger Gemini 2.5 models.

Cheap Requests Can Still Carry Large Working Sets

The context can include:

  • text;
  • images;
  • video;
  • audio;
  • PDFs;
  • conversation history;
  • system instructions;
  • retrieved files;
  • URLs;
  • tool results.

Google Cloud documents approximately 8.4 hours of audio within the 1M-token limit and support for multiple videos per request.

This makes Flash-Lite useful for high-volume multimodal preprocessing.

Examples include:

  • document extraction;
  • transcript summarization;
  • video metadata generation;
  • long-form classification;
  • RAG preprocessing;
  • archive enrichment.

Still, a 1M-token ceiling is not an instruction to fill every request.

Flash-Lite is inexpensive, but sending unnecessary context still increases cost and can reduce relevance.

Use File Search, retrieval, caching, chunking, and selective history where they improve the workload.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

The context can combine text, images, video, audio, PDFs, retrieved content, instructions, history, and tool results.

06 / Thinking budget

Gemini 2.5 Flash-Lite Does Not Think by Default

The most distinctive reasoning behavior of Gemini 2.5 Flash-Lite is its default: if you do not set a thinking budget, the model does not use thinking.

Reasoning Is an Explicit Upgrade Path

Google documents the following thinkingBudget behavior:

  • default: model does not think;
  • thinkingBudget = 0: explicitly disable thinking;
  • fixed budget range: 512 to 24576;
  • thinkingBudget = -1: enable dynamic thinking.

This differs from Gemini 2.5 Flash, which defaults to dynamic thinking.

It also differs from Gemini 3 models, which primarily use thinking_level rather than an explicit token budget.

That makes Flash-Lite useful for applications where the cheapest, fastest behavior should be the default.

For simple classification or extraction, keep thinking disabled.

For ambiguous records or harder subagent tasks, enable a small fixed budget.

For workloads with highly variable complexity, test dynamic thinking.

Thinking tokens are billed at the output-token rate.

Reasoning is off by default
off · default512 · minimum fixed4K · moderate24.6K · maximumdynamic

Routine workload

Keep thinking off for classification, extraction, translation, and routing when evaluation shows reasoning does not improve acceptance.

Ambiguous workload

Enable a fixed or dynamic budget for harder cases before escalating to a more expensive model.

07 / Tools

Search, Maps, Files, Code, URLs, Functions, and Structured Outputs

Gemini 2.5 Flash-Lite is a low-cost model, but it still supports a broad tool set for grounded retrieval, RAG, automation, and structured application workflows.

Lite Can Still Operate Inside an Agent Pipeline

The current Gemini API model page lists:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • context caching.

This makes the model useful for more than static transformations.

A low-cost workflow can:

  1. retrieve from indexed files;
  2. inspect a URL;
  3. execute a calculation;
  4. call an application function;
  5. return a validated structured result.

Flash-Lite does not support the Gemini Live API.

Google Cloud also lists Computer Use as unsupported for this generation.

Those limitations reinforce its strongest role: backend processing, retrieval, transformation, classification, extraction, and narrow tool-based automation rather than rich real-time or GUI-agent experiences.

Low-cost tool surface
  • Google Search

    Ground responses in current web information.

    Supported
  • Google Maps

    Use Maps grounding in supported workflows.

    Supported
  • File Search

    Retrieve relevant content from indexed files.

    Supported
  • Code execution

    Run code for calculations, validation, and transformations.

    Supported
  • Function calling

    Invoke application-defined tools and actions.

    Supported
  • Structured outputs

    Return machine-readable responses that follow an expected schema.

    Supported
  • Computer Use

    Google Cloud lists Computer Use as unsupported for Gemini 2.5 Flash-Lite.

    Not listed
  • Live API

    The base Gemini 2.5 Flash-Lite model does not support Gemini Live API.

    Not listed

08 / Tuning

Gemini 2.5 Flash-Lite Supports Supervised and Preference Tuning

Google Cloud supports both supervised fine-tuning and preference tuning for Gemini 2.5 Flash-Lite, making the inexpensive tier interesting for repetitive specialized workloads.

Customization Can Compound at High Volume

Google Cloud lists:

  • supervised fine-tuning;
  • continuous tuning;
  • preference tuning;
  • tuning checkpoints.

Potential use cases include:

  • domain-specific classification;
  • specialized extraction;
  • stable terminology;
  • organization-specific rewriting;
  • repetitive support tasks;
  • preference-sensitive output style.

Google recommends setting the thinking budget to 0 during supervised fine-tuning for supported Gemini 2.5 thinking models.

That fits Flash-Lite’s core economics.

The goal of tuning is usually to improve task-specific consistency without paying reasoning overhead on every request.

Before tuning, establish a strong prompt-only baseline and measure:

  • task accuracy;
  • schema validity;
  • prompt size;
  • latency;
  • training and maintenance overhead;
  • regression risk;
  • cost per accepted result.
Customize a low-cost production model
  1. 01

    Baseline

    Measure the base model on real production examples and explicit acceptance criteria.

  2. 02

    Tune

    Use supervised or preference tuning when repeated behavior is difficult to stabilize through prompting alone.

  3. 03

    Keep thinking minimal

    Evaluate tuned tasks with thinking disabled when the learned behavior already provides the required specialization.

  4. 04

    Monitor

    Use checkpoints and regression sets to catch quality drift as data and requirements change.

09 / 2.5 Lite vs 2.5 Flash

Gemini 2.5 Flash-Lite vs Gemini 2.5 Flash

Both models share a 1M-token context and much of the same multimodal tool surface, but Flash-Lite optimizes for minimum cost while Flash optimizes for reasoning-heavy price-performance.

Their Default Reasoning Behavior Is the Clearest Difference

Gemini 2.5 Flash-Lite:

  • $0.10 text/image/video input;
  • $0.40 output;
  • thinking off by default;
  • fixed thinking budget starts at 512 tokens.

Gemini 2.5 Flash:

  • $0.30 text/image/video input;
  • $2.50 output;
  • dynamic thinking by default;
  • thinking budget can range from 0 to 24,576 tokens.

That makes regular Flash three times more expensive on text input and more than six times more expensive on output.

The price difference only matters if Flash-Lite can meet the task’s quality threshold.

Use Flash-Lite for narrow, high-volume requests.

Use Flash when ambiguous reasoning, more complex tool orchestration, or difficult multimodal interpretation materially improves completion.

A strong architecture can use both:

  1. run the routine path on Flash-Lite;
  2. validate the result;
  3. increase Flash-Lite thinking for borderline cases;
  4. escalate persistent failures to Flash.
Choose the 2.5 tier
  1. 01

    Start with Flash-Lite

    Use it for cheap classification, extraction, translation, routing, and preprocessing.

  2. 02

    Enable reasoning selectively

    Add a fixed or dynamic Flash-Lite thinking budget only where evaluation shows it improves acceptance.

  3. 03

    Escalate to Flash

    Move difficult reasoning or persistent failures to the stronger Gemini 2.5 Flash tier.

  4. 04

    Measure total economics

    Compare retries, fallbacks, thinking tokens, latency, and accepted-result rate rather than published prices alone.

10 / Lifecycle

Stable in the API, but New Projects Should Use Newer Gemini Models

Gemini 2.5 Flash-Lite is not deprecated in the Gemini API, but Google is increasingly reserving the 2.5 family for existing active users.

Stable Does Not Mean Strategic Default

Google’s current Gemini API documentation states that 2.5 models are not deprecated and will continue to be served until further notice.

At the same time, Google limits access to users who have actively used the 2.5 family in the past.

For new projects, Google recommends newer models such as:

  • Gemini 3.5 Flash-Lite;
  • Gemini 3.8 Flash.

That makes Gemini 2.5 Flash-Lite a legacy production model rather than an obsolete model.

Existing workloads may keep it because:

  • its unit cost is extremely low;
  • behavior is already validated;
  • tuning has been invested in;
  • request volume makes migration economics important.

Google Cloud also lists a platform-specific retirement date of October 20, 2026 for Gemini 2.5 Flash-Lite on the Gemini Enterprise Agent Platform.

That date does not represent a universal Gemini API shutdown; Google explicitly notes that models may remain accessible through the Gemini API after platform retirement.

Stable legacy model
  1. 01

    Not API-deprecated

    Google says stable Gemini 2.5 models will continue to be served through the Gemini API until further notice.

  2. 02

    Legacy access

    Access is increasingly limited to users who have actively used the 2.5 family.

  3. 03

    New-project guidance

    Google recommends Gemini 3.5 Flash-Lite or Gemini 3.8 Flash for new projects.

  4. 04

    Platform retirement differs

    Google Cloud lists October 20, 2026 retirement on Gemini Enterprise Agent Platform, not a universal Gemini API shutdown.

11 / Migration

Migrating From Gemini 2.5 Flash-Lite to Gemini 3.x

The migration is not only about capability: newer Lite models change reasoning controls, tool availability, latency, and cost.

2.5 Flash-Lite Can Still Win on Pure Unit Cost

Google recommends newer models for new projects, but migration economics are workload-dependent.

Gemini 2.5 Flash-Lite standard pricing:

  • $0.10 input;
  • $0.40 output.

Gemini 3.1 Flash-Lite:

  • $0.25 input;
  • $1.50 output.

Gemini 3.5 Flash-Lite:

  • $0.30 input;
  • $2.50 output.

The newer models cost substantially more per output token.

They can justify that premium when stronger reasoning, better agentic performance, newer tools, or fewer failed attempts reduce the total cost of completion.

Reasoning controls also change.

Gemini 2.5 Flash-Lite uses thinkingBudget.

Gemini 3 Lite models use thinking_level.

Do not translate thinkingBudget = 0, 512, or -1 mechanically into a Gemini 3 level.

Replay real workloads and measure:

  • quality;
  • latency;
  • thinking tokens;
  • tool behavior;
  • retries;
  • fallback frequency;
  • total cost.
Move to a newer Lite model
  1. 01

    Map reasoning controls

    Replace explicit 2.5 thinking budgets with evaluated Gemini 3 thinking levels.

  2. 02

    Replay production cases

    Compare extraction, classification, translation, RAG, and multimodal workloads under identical acceptance criteria.

  3. 03

    Measure fewer failures

    A more expensive newer model can still be cheaper if it eliminates enough retries and escalations.

  4. 04

    Keep the cheap path where it wins

    For validated legacy workloads, retain 2.5 Flash-Lite until migration produces a measurable operational benefit.

12 / Evaluation

Gemini 2.5 Flash-Lite Strengths and Limitations

Gemini 2.5 Flash-Lite remains an unusually inexpensive multimodal model with optional reasoning and strong backend tooling, but it now sits on the legacy side of Google’s model portfolio.

Strengths

  • Extremely low API cost

    $0.10 text/image/video input and $0.40 output pricing makes high-volume workloads inexpensive.

  • Thinking off by default

    The model avoids reasoning overhead unless the application explicitly enables a fixed or dynamic thinking budget.

  • 1M-token multimodal context

    Text, images, video, audio, PDFs, retrieval, and tool results can share a large working context.

  • Tuning support

    Google Cloud supports supervised, continuous, and preference tuning plus checkpoints for specialized high-volume behavior.

What to consider

  • Legacy access

    Google is limiting Gemini 2.5 access to existing active users and recommends Gemini 3 for new projects.

  • Weaker agentic ceiling

    Newer Gemini Lite and Flash generations are better suited to difficult coding, computer use, knowledge work, and long-horizon agents.

  • No Computer Use

    Google Cloud lists Computer Use as unsupported for Gemini 2.5 Flash-Lite.

  • No Live API

    Real-time native audio and live conversational workflows require other Gemini model IDs.

Benchmark the cheapest production path

Test Gemini 2.5 Flash-Lite on your real high-volume workload

Replay classification, extraction, translation, routing, document processing, multimodal preprocessing, RAG, and batch jobs to compare pass rate, latency, reasoning overhead, fallback frequency, and cost per accepted result.

Start Free

Gemini 2.5 Flash-Lite is most valuable when a workload can remain on the no-thinking or small-budget path without frequent retries or escalation to a larger model.

Common Questions

What is Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite is Google’s smallest and most cost-efficient Gemini 2.5 model, designed for high-volume classification, simple data extraction, translation, routing, multimodal preprocessing, and extremely low-latency workloads.

What is the Gemini 2.5 Flash-Lite model ID?

The stable model ID is gemini-2.5-flash-lite.

When was Gemini 2.5 Flash-Lite released?

Google introduced Gemini 2.5 Flash-Lite in preview on June 17, 2025 and released the stable generally available model on July 22, 2025.

Is Gemini 2.5 Flash-Lite deprecated?

No. Google’s Gemini API documentation says the stable 2.5 models are not deprecated and will continue to be served until further notice, although access is increasingly limited to existing active users.

Can new users access Gemini 2.5 Flash-Lite?

Google says access to Gemini 2.5 models is being limited to users who have actively used them in the past. For new projects, Google recommends newer Gemini models.

What is the Gemini 2.5 Flash-Lite context window?

Gemini 2.5 Flash-Lite supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 2.5 Flash-Lite support?

The model accepts text, images, video, audio, and PDF input and generates text output.

What is the Gemini 2.5 Flash-Lite knowledge cutoff?

Google documents January 2025 as the knowledge cutoff.

How much does Gemini 2.5 Flash-Lite cost?

Standard paid pricing is $0.10 per 1M text, image, or video input tokens, $0.30 per 1M audio input tokens, and $0.40 per 1M output tokens including thinking.

How much does Gemini 2.5 Flash-Lite context caching cost?

Standard cache reads cost $0.01 per 1M text, image, or video tokens and $0.03 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.

What are Gemini 2.5 Flash-Lite Batch prices?

Batch pricing is $0.05 per 1M text, image, or video input tokens, $0.15 per 1M audio input tokens, and $0.20 per 1M output tokens. Flex uses the same core token rates.

Does Gemini 2.5 Flash-Lite have a free tier?

Yes. Google currently lists free Standard token usage for Gemini 2.5 Flash-Lite in the Gemini Developer API, subject to provider limits and terms.

Does Gemini 2.5 Flash-Lite support thinking?

Yes, but thinking is disabled by default. You can explicitly enable a fixed reasoning budget or dynamic thinking.

What is the Gemini 2.5 Flash-Lite thinking budget?

The fixed thinkingBudget range is 512 to 24,576 tokens. Set thinkingBudget to 0 to disable thinking and -1 to enable dynamic thinking.

Does Gemini 2.5 Flash-Lite think by default?

No. If no thinking budget is supplied, Gemini 2.5 Flash-Lite does not think by default.

How is thinking different between Gemini 2.5 Flash-Lite and Gemini 2.5 Flash?

Flash-Lite defaults to no thinking and supports fixed budgets from 512 to 24,576 tokens. Gemini 2.5 Flash defaults to dynamic thinking and supports a 0 to 24,576-token budget range.

Is Gemini 2.5 Flash-Lite good for classification?

Yes. Google explicitly identifies high-volume classification as one of its primary workloads.

Is Gemini 2.5 Flash-Lite good for extraction and translation?

Yes. Google positions the model for simple data extraction, translation, intelligent routing, and other cost-sensitive high-scale operations.

What tools does Gemini 2.5 Flash-Lite support?

The Gemini API lists Search grounding, Maps grounding, File Search, code execution, function calling, structured outputs, URL context, and context caching.

Does Gemini 2.5 Flash-Lite support Computer Use?

No. Google Cloud currently lists Computer Use as unsupported for Gemini 2.5 Flash-Lite.

Does Gemini 2.5 Flash-Lite support Live API?

No. The base Gemini 2.5 Flash-Lite model does not support Gemini Live API.

Does Gemini 2.5 Flash-Lite support tuning?

Yes on Google Cloud. Google lists supervised fine-tuning, continuous tuning, preference tuning, and tuning checkpoints.

Is Gemini 2.5 Flash-Lite being retired on October 20, 2026?

Google Cloud’s Gemini Enterprise Agent Platform lists October 20, 2026 as the retirement date on that platform. Google’s general Gemini API documentation still says stable Gemini 2.5 models are not deprecated and will continue to be served until further notice.

How is Gemini 2.5 Flash-Lite different from Gemini 3.1 Flash-Lite?

Gemini 2.5 Flash-Lite is much cheaper at $0.10 input and $0.40 output and uses explicit thinkingBudget controls. Gemini 3.1 Flash-Lite costs more but provides a newer model generation and uses Gemini 3 thinking levels.

How is Gemini 2.5 Flash-Lite different from Gemini 3.5 Flash-Lite?

Gemini 2.5 Flash-Lite is optimized for very cheap legacy high-volume processing. Gemini 3.5 Flash-Lite costs more but provides stronger agentic coding, knowledge work, long-context quality, and Computer Use support.

Should I use Gemini 2.5 Flash-Lite for a new application?

For a new application, Google recommends newer Gemini models such as Gemini 3.5 Flash-Lite or Gemini 3.8 Flash. Gemini 2.5 Flash-Lite remains most relevant for established workloads that already depend on its extremely low cost, validated behavior, or tuning investments.

Model information

Last updated

Specifications, pricing, context limits, thinking-budget behavior, multimodal input, tools, tuning support, lifecycle status, access restrictions, and migration guidance on this page are based on official Google Gemini API, Google Cloud, and Google product documentation.

Gemini 2.5 Flash-Lite — Pricing, 1M Context, Thinking Budget & Batch Cost | EidoStack