Google Gemini efficiency model

Gemini 3.1 Flash-Lite

Google’s stable long-term efficiency model for ultra-low-latency, high-volume tasks such as extraction, classification, translation, routing, lightweight agents, and large-scale multimodal processing.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.25
text/image/video · per 1M
Cached input
$0.025
text/image/video · per 1M
Output
$1.50
per 1M tokens

01 / Overview

What Gemini 3.1 Flash-Lite Is

Gemini 3.1 Flash-Lite is Google’s stable efficiency model for high-volume workloads that need low latency, low cost, multimodal input, and enough reasoning and tool use to operate inside production pipelines.

The Cost-Efficiency Tier of Gemini 3.1

Google introduced Gemini 3.1 Flash-Lite in preview on March 3, 2026 and released the stable generally available gemini-3.1-flash-lite endpoint on May 7, 2026.

Google describes it as a cost-efficient model optimized for high-volume agentic tasks, translation, and simple data processing.

Typical workloads include:

  • classification;
  • extraction;
  • translation;
  • routing;
  • simple data processing;
  • document pipelines;
  • multimodal preprocessing;
  • repetitive application workflows.

The model supports up to 1,048,576 input tokens and up to 65,536 output tokens.

Inputs can include text, images, video, audio, and PDFs. Output is text.

Google Cloud lists the stable model as generally available with retirement no earlier than May 7, 2027 on the Gemini Enterprise Agent Platform.

Model profile
Provider
Google
Family
Gemini 3.1 Flash-Lite
Model ID
gemini-3.1-flash-lite
Status
Stable · GA
Knowledge cutoff
Jan 2025
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / Low latency

Ultra-Low Latency Is a Primary Design Goal

Gemini 3.1 Flash-Lite was built for applications where every request must be inexpensive and responsive enough to run at large scale.

Faster Time to First Useful Output

At launch, Google reported that Gemini 3.1 Flash-Lite delivered:

  • 2.5× faster Time to First Answer Token than Gemini 2.5 Flash;
  • 45% higher output speed on Artificial Analysis measurements;
  • similar or better quality for its target model tier.

For high-frequency systems, measure more than an isolated demo response.

Useful metrics include:

  • time to first answer token;
  • total generation time;
  • p50 and p95 latency;
  • throughput under concurrency;
  • retry latency;
  • fallback latency;
  • end-to-end time to an accepted result.

A model can be more valuable because it clears the required quality threshold quickly, even when a larger model scores higher on difficult benchmarks.

Speed at scale
  1. 01

    Fast first token

    Google reported 2.5× faster Time to First Answer Token than Gemini 2.5 Flash at launch.

  2. 02

    Higher output speed

    Google reported a 45% increase in output speed on Artificial Analysis measurements.

  3. 03

    High-frequency execution

    The model is designed for request volumes where small latency improvements compound across many calls.

  4. 04

    Measure full completion

    Include retries, validation, and fallbacks in the real latency of a successful workload.

03 / High-volume workloads

Use the Cheapest Model That Reliably Passes

Gemini 3.1 Flash-Lite is most compelling when a task is repeatable and measurable enough that the low-cost tier can process most requests without escalation.

Separate Routine Work From Difficult Work

A production pipeline can send predictable requests to Flash-Lite and reserve larger models for cases that actually need deeper reasoning.

Good candidates include:

  • intent classification;
  • content tagging;
  • entity extraction;
  • JSON generation;
  • routing;
  • short summarization;
  • normalization;
  • document metadata extraction;
  • data cleanup;
  • simple transformations;
  • lightweight function decisions.

The key metric is not token price alone.

Track:

  • pass rate;
  • schema-valid output rate;
  • retry rate;
  • escalation rate;
  • validation failures;
  • average token usage;
  • blended cost after fallbacks.
Production routing
  1. 01

    Route routine work

    Send predictable, narrow tasks to the low-cost model first.

  2. 02

    Validate

    Use schemas, tests, or business constraints to determine whether the result is acceptable.

  3. 03

    Escalate selectively

    Move difficult or failed cases to a stronger Flash or Pro model.

  4. 04

    Track blended economics

    Include retries and fallback traffic when measuring actual cost per successful task.

04 / Translation & processing

Translation and Simple Data Processing Are Core Use Cases

Google explicitly positions Gemini 3.1 Flash-Lite for translation and simple data processing, making it a natural fit for repetitive language and transformation pipelines.

Scale Language Work Without a Frontier Model

A multilingual application may need to process:

  • UI strings;
  • help-center content;
  • support messages;
  • product descriptions;
  • generated summaries;
  • user-generated content;
  • internal documents;
  • metadata.

Most of these workflows benefit from consistency, latency, and low unit cost more than from maximum reasoning depth.

The same applies to structured processing. Flash-Lite can normalize text, extract fields, classify records, or transform unstructured data before it reaches another system.

For evaluation, measure terminology consistency, preservation of names and numbers, formatting, schema validity, language detection, hallucinated additions, and cost per record.

Language and data pipelines
  1. 01

    Translate

    Process high-volume multilingual text where unit cost and latency matter.

  2. 02

    Normalize

    Convert inconsistent input into predictable application-ready formats.

  3. 03

    Extract

    Pull entities, labels, values, or structured fields from text and documents.

  4. 04

    Validate

    Check terminology, formatting, numeric fidelity, schema validity, and unwanted additions.

05 / Pricing

Gemini 3.1 Flash-Lite Pricing

Gemini 3.1 Flash-Lite costs $0.25 per million text, image, or video input tokens and $1.50 per million output tokens on the standard paid Gemini Developer API tier.

Audio Input Has a Separate Rate

Standard pricing is:

  • $0.25 per 1M text, image, or video input tokens;
  • $0.50 per 1M audio input tokens;
  • $1.50 per 1M output tokens, including thinking;
  • $0.025 per 1M cached text, image, or video tokens;
  • $0.05 per 1M cached audio tokens;
  • $1.00 per 1M cached tokens per hour of storage.

Batch pricing is:

  • $0.125 per 1M text, image, or video input tokens;
  • $0.25 per 1M audio input tokens;
  • $0.75 per 1M output tokens.

Flex uses the same discounted token rates as Batch.

Priority pricing is:

  • $0.45 per 1M text, image, or video input tokens;
  • $0.90 per 1M audio input tokens;
  • $2.70 per 1M output tokens.
Standard and discounted inference

1M tokens · USD

Text / image / video input
$0.25
Audio input
$0.50
Cached text / image / video
$0.025
Output + thinking
$1.50

Example: 20K text input + 4K output

Input cost
$0.0050
Output cost
$0.0060
Estimated total before extra thinking
$0.0110

06 / Multimodal context

A 1M-Token Context Window on the Low-Cost Tier

Gemini 3.1 Flash-Lite supports up to 1,048,576 input tokens and up to 65,536 output tokens despite its latency-first and cost-efficient positioning.

Lite Does Not Mean Short Context

The model can process text, images, video, audio, PDFs, system instructions, history, retrieved content, and tool results.

Google Cloud documents support for up to 3,000 images per prompt, up to 3,000 document pages per file, approximately 45 minutes of video with audio, approximately one hour of video without audio, and approximately 8.4 hours of audio within token limits.

This makes the model useful for:

  • long document extraction;
  • transcript processing;
  • multimodal classification;
  • video and audio preprocessing;
  • batch summarization;
  • high-volume RAG.

The maximum context size is still a ceiling, not a target. Retrieval, File Search, caching, chunking, and selective history can reduce cost and improve relevance.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Text, images, video, audio, documents, instructions, history, retrieved data, and tool results can share the same working context.

07 / Thinking

Minimal Thinking Is the Default

Gemini 3.1 Flash-Lite defaults to minimal thinking, aligning its default behavior with high-throughput and latency-sensitive workloads.

Increase Reasoning Only When the Workload Earns It

Supported levels are:

  • minimal;
  • low;
  • medium;
  • high.

For most requests, minimal behaves similarly to a no-thinking mode, although Google notes that the model can still perform very limited reasoning on complex tasks.

Use minimal for extraction, classification, translation, routing, and simple transformations.

Use low or medium when the same workload benefits from more reasoning without requiring a larger model.

Use high when a difficult subtask still makes economic sense on the Lite tier.

Thinking tokens are billed as output.

Thinking levels
minimal · defaultlowmediumhigh

Routine pipeline

Keep minimal for simple high-volume work when deeper reasoning does not improve acceptance.

Borderline task

Raise reasoning before automatically escalating to a more expensive model.

08 / Tools

Search, Files, Code, URLs, Functions, and Structured Outputs

Gemini 3.1 Flash-Lite supports enough tooling to operate inside low-cost retrieval, transformation, and automation pipelines rather than functioning only as a text completion model.

Tool Use Extends the Lite Tier

Google documents support for:

  • Google Search grounding;
  • File Search;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • context caching;
  • Interactions API;
  • Batch inference;
  • Flex inference;
  • Priority inference.

File Search is especially useful for low-cost RAG and document-processing workflows.

Code execution supports calculations and transformations.

Function calling allows integration with application-defined actions.

Google Cloud currently lists Computer Use and Gemini Live API as unsupported for this model.

Supported production tools
  • Google Search

    Ground supported workflows in current web information.

    Supported
  • File Search

    Retrieve information from indexed files for RAG and document workflows.

    Supported
  • Code execution

    Run code for calculations, transformations, and validation.

    Supported
  • Function calling

    Invoke application-defined tools and actions.

    Supported
  • Structured outputs

    Return schema-constrained machine-readable responses.

    Supported
  • URL context

    Read and reason over supplied URLs.

    Supported
  • Computer Use

    Google Cloud currently lists Computer Use as unsupported.

    Not listed
  • Live API

    Google Cloud currently lists Gemini Live API as unsupported.

    Not listed

09 / Tuning

Gemini 3.1 Flash-Lite Supports Tuning on Google Cloud

Gemini 3.1 Flash-Lite supports supervised fine-tuning and continuous tuning on Google Cloud, giving high-volume applications a path to specialized behavior beyond prompt engineering.

Customization Can Matter More at High Volume

Google Cloud lists:

  • supervised fine-tuning;
  • continuous tuning;
  • tuning checkpoints.

Potential use cases include:

  • domain-specific classification;
  • standardized extraction;
  • organization-specific terminology;
  • stable formatting behavior;
  • repeated transformation tasks.

Tuning should be evaluated against a strong prompt-only baseline.

Measure accuracy, prompt length, latency, regression risk, maintenance overhead, and cost per accepted result.

Customization path
  1. 01

    Baseline

    Measure the untuned model with real production prompts and acceptance criteria.

  2. 02

    Tune

    Use supervised or continuous tuning when narrow behavior is difficult to maintain through prompts alone.

  3. 03

    Compare

    Evaluate quality, prompt size, latency, and cost under identical workloads.

  4. 04

    Monitor

    Keep regression tests and checkpoints around customized behavior.

10 / 3.1 Lite vs 3.5 Lite

Gemini 3.1 Flash-Lite vs Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite is the stronger newer Lite model for agentic coding, computer use, knowledge work, and harder long-context tasks, while Gemini 3.1 Flash-Lite remains the cheaper stable baseline for routine high-volume processing.

The Upgrade Is a Price-Performance Decision

Gemini 3.1 Flash-Lite standard pricing:

  • $0.25 text/image/video input;
  • $1.50 output.

Gemini 3.5 Flash-Lite standard pricing:

  • $0.30 input;
  • $2.50 output.

The newer model costs more, especially on output.

Google’s 3.5 evaluations show stronger performance in agentic coding, computer use, knowledge work, and long-context execution. It also supports Computer Use, which 3.1 Flash-Lite does not.

Google still recommends 3.1 Flash-Lite as a stable long-term model for low-cost, high-volume tasks that do not require advanced reasoning depth.

A useful routing strategy is:

  1. keep extraction, translation, classification, and simple transformations on 3.1;
  2. move harder subagents and agentic workloads to 3.5 Lite;
  3. measure retries and fallback frequency;
  4. optimize for cost per accepted result.
Upgrade or keep the cheaper baseline
  1. 01

    Keep routine workloads on 3.1

    Use the lower output price where the model already meets acceptance criteria.

  2. 02

    Move difficult agents to 3.5

    Use the newer model for coding, Computer Use, harder knowledge work, and persistent failures.

  3. 03

    Measure fallbacks

    A cheaper model loses its advantage if too many requests must be repeated on the stronger tier.

  4. 04

    Optimize the mix

    Route by workload economics rather than using one model for every request.

11 / Evaluation

Gemini 3.1 Flash-Lite Strengths and Limitations

Gemini 3.1 Flash-Lite remains a strong stable efficiency baseline: inexpensive, fast, multimodal, tool-capable, and tunable, but intentionally less capable than newer Flash tiers on difficult agentic work.

Strengths

  • Very low API cost

    $0.25 text/image/video input and $1.50 output pricing supports large request volumes.

  • Low-latency design

    Google reported 2.5× faster first-answer-token latency and 45% higher output speed than Gemini 2.5 Flash at launch.

  • 1M-token multimodal context

    Text, images, video, audio, and documents fit within a large context even on the Lite tier.

  • Tuning support

    Google Cloud supports supervised fine-tuning, continuous tuning, and tuning checkpoints.

What to consider

  • Not the strongest Lite generation

    Gemini 3.5 Flash-Lite improves agentic coding, computer use, knowledge work, and hard long-context execution.

  • No Computer Use

    Google Cloud currently lists Computer Use as unsupported for Gemini 3.1 Flash-Lite.

  • No Live API

    The model is built for request-response and batch-style workloads rather than native Gemini Live interactions.

  • Routine-task optimization

    Difficult reasoning and long-horizon agents may require retries or escalation to a stronger model.

Measure the economics of scale

Test Gemini 3.1 Flash-Lite on your production workload

Replay extraction, classification, routing, translation, multimodal parsing, batch jobs, and lightweight agent tasks to compare accepted-result rate, latency, thinking tokens, fallback frequency, and cost per completed request.

Start Free

Gemini 3.1 Flash-Lite is most valuable when a simple, repeatable workload can stay on the low-cost path without frequent retries or escalation to a larger model.

Common Questions

What is Gemini 3.1 Flash-Lite?

Gemini 3.1 Flash-Lite is Google’s stable cost-efficiency model for high-volume agentic subtasks, translation, classification, extraction, routing, and simple data processing.

What is the Gemini 3.1 Flash-Lite model ID?

The stable model ID is gemini-3.1-flash-lite.

When was Gemini 3.1 Flash-Lite released?

Google introduced Gemini 3.1 Flash-Lite in preview on March 3, 2026 and released the stable generally available endpoint on May 7, 2026.

Is Gemini 3.1 Flash-Lite stable?

Yes. Google released gemini-3.1-flash-lite as generally available on May 7, 2026.

What happened to gemini-3.1-flash-lite-preview?

Google shut down gemini-3.1-flash-lite-preview on May 25, 2026. The stable replacement is gemini-3.1-flash-lite.

What is the Gemini 3.1 Flash-Lite context window?

Gemini 3.1 Flash-Lite supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.1 Flash-Lite support?

The model supports text, images, video, audio, and document input such as PDFs and produces text output.

What is the Gemini 3.1 Flash-Lite knowledge cutoff?

Google documents January 2025 as the Gemini 3.1 Flash-Lite knowledge cutoff.

How much does Gemini 3.1 Flash-Lite cost?

Standard paid pricing is $0.25 per 1M text, image, or video input tokens, $0.50 per 1M audio input tokens, and $1.50 per 1M output tokens including thinking.

How much does Gemini 3.1 Flash-Lite context caching cost?

Standard cache reads cost $0.025 per 1M text, image, or video tokens and $0.05 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.

What are the Batch prices for Gemini 3.1 Flash-Lite?

Batch pricing is $0.125 per 1M text, image, or video input tokens, $0.25 per 1M audio input tokens, and $0.75 per 1M output tokens. Flex uses the same token rates.

Does Gemini 3.1 Flash-Lite have a free tier?

Yes. Google currently lists free standard token usage for Gemini 3.1 Flash-Lite in the Gemini Developer API, subject to provider limits and terms.

Does Gemini 3.1 Flash-Lite support thinking?

Yes. It supports minimal, low, medium, and high thinking levels.

What is the default thinking level of Gemini 3.1 Flash-Lite?

Minimal is the default thinking level, prioritizing low latency and high throughput.

How fast is Gemini 3.1 Flash-Lite?

At launch, Google reported 2.5× faster Time to First Answer Token and 45% higher output speed than Gemini 2.5 Flash.

Is Gemini 3.1 Flash-Lite good for translation?

Yes. Google explicitly lists translation as one of the model’s target workloads.

Is Gemini 3.1 Flash-Lite good for extraction and classification?

Yes. Its low price, minimal-default thinking, structured outputs, multimodal input, and large context make it suitable for high-volume extraction, classification, routing, and transformation.

Does Gemini 3.1 Flash-Lite support File Search?

Yes. Google’s Gemini API File Search documentation lists Gemini 3.1 Flash-Lite as supported.

Does Gemini 3.1 Flash-Lite support Computer Use?

No. Google Cloud currently lists Computer Use as unsupported for Gemini 3.1 Flash-Lite.

What tools does Gemini 3.1 Flash-Lite support?

Google documents support for Search grounding, File Search, code execution, function calling, structured outputs, URL context, and context caching. Batch, Flex, and Priority serving options are also available.

Does Gemini 3.1 Flash-Lite support tuning?

Yes on Google Cloud. The current model specification lists supervised fine-tuning, continuous tuning, and tuning checkpoints.

Does Gemini 3.1 Flash-Lite support the Live API?

No. Google Cloud currently lists Gemini Live API as unsupported for this model.

How is Gemini 3.1 Flash-Lite different from Gemini 3.5 Flash-Lite?

Gemini 3.1 Flash-Lite is the cheaper stable efficiency baseline at $0.25 input and $1.50 output. Gemini 3.5 Flash-Lite costs more but improves agentic coding, computer use, knowledge work, and difficult long-context execution.

Should I use Gemini 3.1 Flash-Lite for a new application?

It remains a strong choice for applications dominated by translation, extraction, classification, routing, preprocessing, and other high-volume cost-sensitive tasks. For harder subagents, coding, or computer-use workflows, compare it directly with Gemini 3.5 Flash-Lite and larger Flash models.

Model information

Last updated

Specifications, pricing, release status, context limits, thinking behavior, performance positioning, multimodal inputs, tools, tuning, serving tiers, lifecycle, and migration guidance on this page are based on official Google Gemini API, Google Cloud, and Google product documentation.

Gemini 3.1 Flash-Lite — Pricing, 1M Context, Low Latency & Tuning | EidoStack