Google Gemini efficiency model
Gemini 3.1 Flash-Lite
Google’s stable long-term efficiency model for ultra-low-latency, high-volume tasks such as extraction, classification, translation, routing, lightweight agents, and large-scale multimodal processing.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $0.25
- text/image/video · per 1M
- Cached input
- $0.025
- text/image/video · per 1M
- Output
- $1.50
- per 1M tokens
01 / Overview
What Gemini 3.1 Flash-Lite Is
Gemini 3.1 Flash-Lite is Google’s stable efficiency model for high-volume workloads that need low latency, low cost, multimodal input, and enough reasoning and tool use to operate inside production pipelines.
The Cost-Efficiency Tier of Gemini 3.1
Google introduced Gemini 3.1 Flash-Lite in preview on March 3, 2026 and released the stable generally available gemini-3.1-flash-lite endpoint on May 7, 2026.
Google describes it as a cost-efficient model optimized for high-volume agentic tasks, translation, and simple data processing.
Typical workloads include:
- classification;
- extraction;
- translation;
- routing;
- simple data processing;
- document pipelines;
- multimodal preprocessing;
- repetitive application workflows.
The model supports up to 1,048,576 input tokens and up to 65,536 output tokens.
Inputs can include text, images, video, audio, and PDFs. Output is text.
Google Cloud lists the stable model as generally available with retirement no earlier than May 7, 2027 on the Gemini Enterprise Agent Platform.
- Provider
- Family
- Gemini 3.1 Flash-Lite
- Model ID
- gemini-3.1-flash-lite
- Status
- Stable · GA
- Knowledge cutoff
- Jan 2025
- Input
- Text, Image, Video, Audio, PDF
- Output
- Text
02 / Low latency
Ultra-Low Latency Is a Primary Design Goal
Gemini 3.1 Flash-Lite was built for applications where every request must be inexpensive and responsive enough to run at large scale.
Faster Time to First Useful Output
At launch, Google reported that Gemini 3.1 Flash-Lite delivered:
- 2.5× faster Time to First Answer Token than Gemini 2.5 Flash;
- 45% higher output speed on Artificial Analysis measurements;
- similar or better quality for its target model tier.
For high-frequency systems, measure more than an isolated demo response.
Useful metrics include:
- time to first answer token;
- total generation time;
- p50 and p95 latency;
- throughput under concurrency;
- retry latency;
- fallback latency;
- end-to-end time to an accepted result.
A model can be more valuable because it clears the required quality threshold quickly, even when a larger model scores higher on difficult benchmarks.
- 01
Fast first token
Google reported 2.5× faster Time to First Answer Token than Gemini 2.5 Flash at launch.
- 02
Higher output speed
Google reported a 45% increase in output speed on Artificial Analysis measurements.
- 03
High-frequency execution
The model is designed for request volumes where small latency improvements compound across many calls.
- 04
Measure full completion
Include retries, validation, and fallbacks in the real latency of a successful workload.
03 / High-volume workloads
Use the Cheapest Model That Reliably Passes
Gemini 3.1 Flash-Lite is most compelling when a task is repeatable and measurable enough that the low-cost tier can process most requests without escalation.
Separate Routine Work From Difficult Work
A production pipeline can send predictable requests to Flash-Lite and reserve larger models for cases that actually need deeper reasoning.
Good candidates include:
- intent classification;
- content tagging;
- entity extraction;
- JSON generation;
- routing;
- short summarization;
- normalization;
- document metadata extraction;
- data cleanup;
- simple transformations;
- lightweight function decisions.
The key metric is not token price alone.
Track:
- pass rate;
- schema-valid output rate;
- retry rate;
- escalation rate;
- validation failures;
- average token usage;
- blended cost after fallbacks.
- 01
Route routine work
Send predictable, narrow tasks to the low-cost model first.
- 02
Validate
Use schemas, tests, or business constraints to determine whether the result is acceptable.
- 03
Escalate selectively
Move difficult or failed cases to a stronger Flash or Pro model.
- 04
Track blended economics
Include retries and fallback traffic when measuring actual cost per successful task.
04 / Translation & processing
Translation and Simple Data Processing Are Core Use Cases
Google explicitly positions Gemini 3.1 Flash-Lite for translation and simple data processing, making it a natural fit for repetitive language and transformation pipelines.
Scale Language Work Without a Frontier Model
A multilingual application may need to process:
- UI strings;
- help-center content;
- support messages;
- product descriptions;
- generated summaries;
- user-generated content;
- internal documents;
- metadata.
Most of these workflows benefit from consistency, latency, and low unit cost more than from maximum reasoning depth.
The same applies to structured processing. Flash-Lite can normalize text, extract fields, classify records, or transform unstructured data before it reaches another system.
For evaluation, measure terminology consistency, preservation of names and numbers, formatting, schema validity, language detection, hallucinated additions, and cost per record.
- 01
Translate
Process high-volume multilingual text where unit cost and latency matter.
- 02
Normalize
Convert inconsistent input into predictable application-ready formats.
- 03
Extract
Pull entities, labels, values, or structured fields from text and documents.
- 04
Validate
Check terminology, formatting, numeric fidelity, schema validity, and unwanted additions.
05 / Pricing
Gemini 3.1 Flash-Lite Pricing
Gemini 3.1 Flash-Lite costs $0.25 per million text, image, or video input tokens and $1.50 per million output tokens on the standard paid Gemini Developer API tier.
Audio Input Has a Separate Rate
Standard pricing is:
- $0.25 per 1M text, image, or video input tokens;
- $0.50 per 1M audio input tokens;
- $1.50 per 1M output tokens, including thinking;
- $0.025 per 1M cached text, image, or video tokens;
- $0.05 per 1M cached audio tokens;
- $1.00 per 1M cached tokens per hour of storage.
Batch pricing is:
- $0.125 per 1M text, image, or video input tokens;
- $0.25 per 1M audio input tokens;
- $0.75 per 1M output tokens.
Flex uses the same discounted token rates as Batch.
Priority pricing is:
- $0.45 per 1M text, image, or video input tokens;
- $0.90 per 1M audio input tokens;
- $2.70 per 1M output tokens.
1M tokens · USD
- Text / image / video input
- $0.25
- Audio input
- $0.50
- Cached text / image / video
- $0.025
- Output + thinking
- $1.50
Example: 20K text input + 4K output
- Input cost
- $0.0050
- Output cost
- $0.0060
- Estimated total before extra thinking
- $0.0110
06 / Multimodal context
A 1M-Token Context Window on the Low-Cost Tier
Gemini 3.1 Flash-Lite supports up to 1,048,576 input tokens and up to 65,536 output tokens despite its latency-first and cost-efficient positioning.
Lite Does Not Mean Short Context
The model can process text, images, video, audio, PDFs, system instructions, history, retrieved content, and tool results.
Google Cloud documents support for up to 3,000 images per prompt, up to 3,000 document pages per file, approximately 45 minutes of video with audio, approximately one hour of video without audio, and approximately 8.4 hours of audio within token limits.
This makes the model useful for:
- long document extraction;
- transcript processing;
- multimodal classification;
- video and audio preprocessing;
- batch summarization;
- high-volume RAG.
The maximum context size is still a ceiling, not a target. Retrieval, File Search, caching, chunking, and selective history can reduce cost and improve relevance.
Input context
1,048,576
Max output
65,536
Text, images, video, audio, documents, instructions, history, retrieved data, and tool results can share the same working context.
07 / Thinking
Minimal Thinking Is the Default
Gemini 3.1 Flash-Lite defaults to minimal thinking, aligning its default behavior with high-throughput and latency-sensitive workloads.
Increase Reasoning Only When the Workload Earns It
Supported levels are:
minimal;low;medium;high.
For most requests, minimal behaves similarly to a no-thinking mode, although Google notes that the model can still perform very limited reasoning on complex tasks.
Use minimal for extraction, classification, translation, routing, and simple transformations.
Use low or medium when the same workload benefits from more reasoning without requiring a larger model.
Use high when a difficult subtask still makes economic sense on the Lite tier.
Thinking tokens are billed as output.
Routine pipeline
Keep minimal for simple high-volume work when deeper reasoning does not improve acceptance.
Borderline task
Raise reasoning before automatically escalating to a more expensive model.
08 / Tools
Search, Files, Code, URLs, Functions, and Structured Outputs
Gemini 3.1 Flash-Lite supports enough tooling to operate inside low-cost retrieval, transformation, and automation pipelines rather than functioning only as a text completion model.
Tool Use Extends the Lite Tier
Google documents support for:
- Google Search grounding;
- File Search;
- code execution;
- function calling;
- structured outputs;
- URL context;
- context caching;
- Interactions API;
- Batch inference;
- Flex inference;
- Priority inference.
File Search is especially useful for low-cost RAG and document-processing workflows.
Code execution supports calculations and transformations.
Function calling allows integration with application-defined actions.
Google Cloud currently lists Computer Use and Gemini Live API as unsupported for this model.
- Supported
Google Search
Ground supported workflows in current web information.
- Supported
File Search
Retrieve information from indexed files for RAG and document workflows.
- Supported
Code execution
Run code for calculations, transformations, and validation.
- Supported
Function calling
Invoke application-defined tools and actions.
- Supported
Structured outputs
Return schema-constrained machine-readable responses.
- Supported
URL context
Read and reason over supplied URLs.
- Not listed
Computer Use
Google Cloud currently lists Computer Use as unsupported.
- Not listed
Live API
Google Cloud currently lists Gemini Live API as unsupported.
09 / Tuning
Gemini 3.1 Flash-Lite Supports Tuning on Google Cloud
Gemini 3.1 Flash-Lite supports supervised fine-tuning and continuous tuning on Google Cloud, giving high-volume applications a path to specialized behavior beyond prompt engineering.
Customization Can Matter More at High Volume
Google Cloud lists:
- supervised fine-tuning;
- continuous tuning;
- tuning checkpoints.
Potential use cases include:
- domain-specific classification;
- standardized extraction;
- organization-specific terminology;
- stable formatting behavior;
- repeated transformation tasks.
Tuning should be evaluated against a strong prompt-only baseline.
Measure accuracy, prompt length, latency, regression risk, maintenance overhead, and cost per accepted result.
- 01
Baseline
Measure the untuned model with real production prompts and acceptance criteria.
- 02
Tune
Use supervised or continuous tuning when narrow behavior is difficult to maintain through prompts alone.
- 03
Compare
Evaluate quality, prompt size, latency, and cost under identical workloads.
- 04
Monitor
Keep regression tests and checkpoints around customized behavior.
10 / 3.1 Lite vs 3.5 Lite
Gemini 3.1 Flash-Lite vs Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is the stronger newer Lite model for agentic coding, computer use, knowledge work, and harder long-context tasks, while Gemini 3.1 Flash-Lite remains the cheaper stable baseline for routine high-volume processing.
The Upgrade Is a Price-Performance Decision
Gemini 3.1 Flash-Lite standard pricing:
- $0.25 text/image/video input;
- $1.50 output.
Gemini 3.5 Flash-Lite standard pricing:
- $0.30 input;
- $2.50 output.
The newer model costs more, especially on output.
Google’s 3.5 evaluations show stronger performance in agentic coding, computer use, knowledge work, and long-context execution. It also supports Computer Use, which 3.1 Flash-Lite does not.
Google still recommends 3.1 Flash-Lite as a stable long-term model for low-cost, high-volume tasks that do not require advanced reasoning depth.
A useful routing strategy is:
- keep extraction, translation, classification, and simple transformations on 3.1;
- move harder subagents and agentic workloads to 3.5 Lite;
- measure retries and fallback frequency;
- optimize for cost per accepted result.
- 01
Keep routine workloads on 3.1
Use the lower output price where the model already meets acceptance criteria.
- 02
Move difficult agents to 3.5
Use the newer model for coding, Computer Use, harder knowledge work, and persistent failures.
- 03
Measure fallbacks
A cheaper model loses its advantage if too many requests must be repeated on the stronger tier.
- 04
Optimize the mix
Route by workload economics rather than using one model for every request.
11 / Evaluation
Gemini 3.1 Flash-Lite Strengths and Limitations
Gemini 3.1 Flash-Lite remains a strong stable efficiency baseline: inexpensive, fast, multimodal, tool-capable, and tunable, but intentionally less capable than newer Flash tiers on difficult agentic work.
Strengths
Very low API cost
$0.25 text/image/video input and $1.50 output pricing supports large request volumes.
Low-latency design
Google reported 2.5× faster first-answer-token latency and 45% higher output speed than Gemini 2.5 Flash at launch.
1M-token multimodal context
Text, images, video, audio, and documents fit within a large context even on the Lite tier.
Tuning support
Google Cloud supports supervised fine-tuning, continuous tuning, and tuning checkpoints.
What to consider
Not the strongest Lite generation
Gemini 3.5 Flash-Lite improves agentic coding, computer use, knowledge work, and hard long-context execution.
No Computer Use
Google Cloud currently lists Computer Use as unsupported for Gemini 3.1 Flash-Lite.
No Live API
The model is built for request-response and batch-style workloads rather than native Gemini Live interactions.
Routine-task optimization
Difficult reasoning and long-horizon agents may require retries or escalation to a stronger model.
Measure the economics of scale
Test Gemini 3.1 Flash-Lite on your production workload
Replay extraction, classification, routing, translation, multimodal parsing, batch jobs, and lightweight agent tasks to compare accepted-result rate, latency, thinking tokens, fallback frequency, and cost per completed request.
Start FreeGemini 3.1 Flash-Lite is most valuable when a simple, repeatable workload can stay on the low-cost path without frequent retries or escalation to a larger model.
Common Questions
What is Gemini 3.1 Flash-Lite?
Gemini 3.1 Flash-Lite is Google’s stable cost-efficiency model for high-volume agentic subtasks, translation, classification, extraction, routing, and simple data processing.
What is the Gemini 3.1 Flash-Lite model ID?
The stable model ID is gemini-3.1-flash-lite.
When was Gemini 3.1 Flash-Lite released?
Google introduced Gemini 3.1 Flash-Lite in preview on March 3, 2026 and released the stable generally available endpoint on May 7, 2026.
Is Gemini 3.1 Flash-Lite stable?
Yes. Google released gemini-3.1-flash-lite as generally available on May 7, 2026.
What happened to gemini-3.1-flash-lite-preview?
Google shut down gemini-3.1-flash-lite-preview on May 25, 2026. The stable replacement is gemini-3.1-flash-lite.
What is the Gemini 3.1 Flash-Lite context window?
Gemini 3.1 Flash-Lite supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3.1 Flash-Lite support?
The model supports text, images, video, audio, and document input such as PDFs and produces text output.
What is the Gemini 3.1 Flash-Lite knowledge cutoff?
Google documents January 2025 as the Gemini 3.1 Flash-Lite knowledge cutoff.
How much does Gemini 3.1 Flash-Lite cost?
Standard paid pricing is $0.25 per 1M text, image, or video input tokens, $0.50 per 1M audio input tokens, and $1.50 per 1M output tokens including thinking.
How much does Gemini 3.1 Flash-Lite context caching cost?
Standard cache reads cost $0.025 per 1M text, image, or video tokens and $0.05 per 1M audio tokens, plus $1 per 1M cached tokens per hour of storage.
What are the Batch prices for Gemini 3.1 Flash-Lite?
Batch pricing is $0.125 per 1M text, image, or video input tokens, $0.25 per 1M audio input tokens, and $0.75 per 1M output tokens. Flex uses the same token rates.
Does Gemini 3.1 Flash-Lite have a free tier?
Yes. Google currently lists free standard token usage for Gemini 3.1 Flash-Lite in the Gemini Developer API, subject to provider limits and terms.
Does Gemini 3.1 Flash-Lite support thinking?
Yes. It supports minimal, low, medium, and high thinking levels.
What is the default thinking level of Gemini 3.1 Flash-Lite?
Minimal is the default thinking level, prioritizing low latency and high throughput.
How fast is Gemini 3.1 Flash-Lite?
At launch, Google reported 2.5× faster Time to First Answer Token and 45% higher output speed than Gemini 2.5 Flash.
Is Gemini 3.1 Flash-Lite good for translation?
Yes. Google explicitly lists translation as one of the model’s target workloads.
Is Gemini 3.1 Flash-Lite good for extraction and classification?
Yes. Its low price, minimal-default thinking, structured outputs, multimodal input, and large context make it suitable for high-volume extraction, classification, routing, and transformation.
Does Gemini 3.1 Flash-Lite support File Search?
Yes. Google’s Gemini API File Search documentation lists Gemini 3.1 Flash-Lite as supported.
Does Gemini 3.1 Flash-Lite support Computer Use?
No. Google Cloud currently lists Computer Use as unsupported for Gemini 3.1 Flash-Lite.
What tools does Gemini 3.1 Flash-Lite support?
Google documents support for Search grounding, File Search, code execution, function calling, structured outputs, URL context, and context caching. Batch, Flex, and Priority serving options are also available.
Does Gemini 3.1 Flash-Lite support tuning?
Yes on Google Cloud. The current model specification lists supervised fine-tuning, continuous tuning, and tuning checkpoints.
Does Gemini 3.1 Flash-Lite support the Live API?
No. Google Cloud currently lists Gemini Live API as unsupported for this model.
How is Gemini 3.1 Flash-Lite different from Gemini 3.5 Flash-Lite?
Gemini 3.1 Flash-Lite is the cheaper stable efficiency baseline at $0.25 input and $1.50 output. Gemini 3.5 Flash-Lite costs more but improves agentic coding, computer use, knowledge work, and difficult long-context execution.
Should I use Gemini 3.1 Flash-Lite for a new application?
It remains a strong choice for applications dominated by translation, extraction, classification, routing, preprocessing, and other high-volume cost-sensitive tasks. For harder subagents, coding, or computer-use workflows, compare it directly with Gemini 3.5 Flash-Lite and larger Flash models.
Model information
Last updated
Specifications, pricing, release status, context limits, thinking behavior, performance positioning, multimodal inputs, tools, tuning, serving tiers, lifecycle, and migration guidance on this page are based on official Google Gemini API, Google Cloud, and Google product documentation.