Google Gemini efficiency model

Gemini 3.5 Flash-Lite

Google’s fastest and most cost-effective 3.5-class model for high-throughput subagents, document parsing, extraction, translation, multimodal processing, and latency-sensitive production workloads.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.30
per 1M tokens
Cached input
$0.03
per 1M tokens
Output
$2.50
per 1M tokens

01 / Overview

What Gemini 3.5 Flash-Lite Is

Gemini 3.5 Flash-Lite is Google’s fastest and most cost-effective model in the Gemini 3.5 generation, optimized for high-throughput execution, low latency, and production workloads where inference cost is a primary constraint.

Flash-Lite Is an Efficiency Tier, Not a Text-Only Utility Model

Google released Gemini 3.5 Flash-Lite as generally available on July 21, 2026.

Its positioning is deliberately different from Gemini 3.5 Flash.

Gemini 3.5 Flash targets frontier-level coding and agentic work. Flash-Lite targets the large volume of requests where a smaller, faster model can meet the acceptance threshold at much lower cost.

Typical workloads include:

  • subagent tasks;
  • classification;
  • routing;
  • document parsing;
  • structured extraction;
  • translation;
  • high-volume multimodal processing;
  • lightweight coding;
  • repetitive tool-based automation.

The model is still natively multimodal.

It accepts text, images, video, audio, and PDFs, supports a 1,048,576-token input window, and can generate up to 65,536 text output tokens.

It also supports thinking, function calling, Search grounding, Maps grounding, File Search, code execution, structured outputs, URL context, caching, and Computer Use in preview.

  • Stable model ID: gemini-3.5-flash-lite.
  • GA since July 21, 2026.
  • 1,048,576 input tokens.
  • 65,536 maximum output tokens.
  • Text, image, video, audio, and PDF input.
  • Text output.
  • Minimal thinking by default.
  • Retirement on Google Cloud no earlier than July 21, 2027.
Model profile
Provider
Google
Family
Gemini 3.5 Flash-Lite
Model ID
gemini-3.5-flash-lite
Status
Stable · GA
Knowledge cutoff
Mar 2026*
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / High throughput

Built for High-Throughput, Latency-Sensitive Production

Gemini 3.5 Flash-Lite is designed for workloads where the application may make thousands or millions of model calls and where every extra millisecond or output token compounds into infrastructure cost.

The Unit of Optimization Is Cost per Accepted Result

DeepMind describes Gemini 3.5 Flash-Lite as the fastest and most cost-effective 3.5-class model.

It reports roughly 350 output tokens per second according to the Artificial Analysis Index.

Raw throughput matters for interactive products, but the production metric should be broader.

A cheap model can become expensive if it:

  • requires frequent retries;
  • produces invalid structured output;
  • misroutes requests;
  • needs premium-model fallback too often;
  • calls unnecessary tools;
  • generates excessive reasoning.

For high-volume applications, track:

  • p50 and p95 latency;
  • accepted-result rate;
  • retry rate;
  • fallback rate;
  • output tokens;
  • thinking tokens;
  • tool calls;
  • cost per completed task.

Flash-Lite is strongest when the quality threshold is clear and the model can meet it consistently without escalating most requests.

Throughput-first workloads
  1. 01

    Classification

    Categorize large request volumes where deterministic labels matter more than deep open-ended reasoning.

  2. 02

    Routing

    Choose workflows, tools, queues, or stronger downstream models at low per-request cost.

  3. 03

    Transformation

    Normalize, rewrite, tag, summarize, or convert content into application-ready formats.

  4. 04

    Batch processing

    Use discounted Batch or Flex inference when workloads are asynchronous and throughput matters more than immediate response.

03 / Subagents

A Strong Fit for Subagents Inside Larger Systems

Google explicitly positions Gemini 3.5 Flash-Lite for high-volume subagent tasks, where a larger coordinator can delegate narrow work to a cheaper specialist.

Not Every Agent Step Needs the Strongest Model

A complex agent can contain many kinds of operations.

Some require deep planning. Others are narrow and repeatable:

  • inspect one document;
  • extract a field;
  • classify a result;
  • summarize a tool response;
  • search for a specific fact;
  • validate a structured object;
  • translate text;
  • choose among known actions.

Sending every subtask to an expensive frontier model can waste budget.

A hierarchical system can reserve stronger models for planning and difficult decisions while using Flash-Lite for the high-volume leaves of the task tree.

This architecture works only if routing is measured.

If cheap subagents fail frequently and force the coordinator to redo the task, the apparent savings disappear.

Measure subagent success independently from the overall workflow and track how often results are rejected or escalated.

Low-cost agent specialization
  1. 01

    Delegate

    Send narrow, well-specified subtasks to Flash-Lite instead of consuming premium-model capacity.

  2. 02

    Validate

    Check structured outputs, extraction accuracy, or task-specific constraints before accepting the result.

  3. 03

    Escalate

    Route difficult or failed cases to Gemini Flash, Pro, or another stronger model.

  4. 04

    Measure blended cost

    Include retries and escalations when calculating the economics of a multi-model agent tree.

04 / Documents & extraction

Document Parsing and Extraction at Scale

Document processing is one of the model’s explicit target workloads: Gemini 3.5 Flash-Lite can combine long context, multimodal input, structured outputs, and low token pricing for large extraction pipelines.

PDFs Are Only One Part of the Pipeline

Real document workflows may include:

  • PDFs;
  • scanned images;
  • receipts;
  • charts;
  • forms;
  • audio;
  • video;
  • text attachments;
  • metadata;
  • business rules.

Gemini 3.5 Flash-Lite can process these inputs natively and return structured text suitable for downstream systems.

DeepMind highlights document processing, search, translation, classification, and agentic workflows as intended uses.

This makes the model useful for:

  • receipt and invoice extraction;
  • form processing;
  • contract metadata;
  • document classification;
  • search enrichment;
  • content moderation support;
  • dataset labeling;
  • large-scale summarization.

For extraction tasks, accuracy is not the only metric.

Measure field-level precision and recall, JSON validity, missing-value behavior, latency, token usage, and exception rate.

A model that costs less per token but causes more manual review may be more expensive operationally.

Document pipeline
  1. 01

    Ingest

    Accept PDFs, text, images, audio, or video inside the same multimodal workflow.

  2. 02

    Extract

    Identify fields, entities, categories, summaries, or other application-specific information.

  3. 03

    Structure

    Return schema-constrained output for databases, queues, APIs, and automation.

  4. 04

    Validate

    Measure field accuracy, malformed-output rate, missing data, and downstream manual-review cost.

05 / Pricing

Gemini 3.5 Flash-Lite Pricing

Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens on the standard paid Gemini Developer API tier.

Batch and Flex Cut Token Rates in Half

Google currently lists Standard pricing at:

  • $0.30 per 1M input tokens for text, image, video, and audio.
  • $2.50 per 1M output tokens, including thinking tokens.
  • $0.03 per 1M cached-context tokens.
  • $1.00 per 1M cached tokens per hour of cache storage.

Batch pricing is:

  • $0.15 per 1M input tokens.
  • $1.25 per 1M output tokens.
  • $0.02 per 1M cached-context tokens.

Flex inference uses the same $0.15 input / $1.25 output rates.

Priority inference is more expensive:

  • $0.54 per 1M input tokens.
  • $4.50 per 1M output tokens.
  • $0.05 per 1M cached-context tokens.

The standard Gemini Developer API also lists free token usage, subject to provider limits and terms.

Because thinking tokens are included in output billing, even an efficiency model should be evaluated using actual API usage rather than visible response length.

Standard and discounted inference

1M tokens · USD

Input
$0.30
Cached input
$0.03
Output + thinking
$2.50

Example: 20K input + 4K output

Input cost
$0.0060
Output cost
$0.0100
Estimated total before extra thinking
$0.0160

06 / Multimodal context

A 1M-Token Window on the Lite Tier

Gemini 3.5 Flash-Lite keeps the same 1,048,576-token input limit and 65,536-token maximum output as larger Gemini Flash models.

Lite Does Not Mean Small Context

The context window can contain:

  • text;
  • images;
  • video;
  • audio;
  • PDFs;
  • system instructions;
  • conversation history;
  • retrieved information;
  • tool results.

This is useful when a low-cost model must process large source material rather than only short classification prompts.

Examples include:

  • long PDFs;
  • multi-document extraction;
  • large transcript processing;
  • video analysis;
  • document search;
  • batch summarization;
  • lightweight repository inspection.

However, the model remains optimized for speed and cost rather than maximum reasoning depth.

A million-token context window does not guarantee that every difficult long-context problem should be assigned to Flash-Lite.

Test retrieval accuracy, needle-in-context performance, output correctness, and cost before using extreme prompt sizes in production.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Flash-Lite retains a 1M-token input window even though it is optimized for lower latency and cost rather than maximum reasoning depth.

07 / Thinking

Minimal Thinking by Default

Gemini 3.5 Flash-Lite defaults to minimal thinking, making latency and throughput the default behavior rather than requiring developers to explicitly disable deeper reasoning.

Increase Reasoning Only When the Workload Earns It

Supported thinking levels are:

  • minimal;
  • low;
  • medium;
  • high.

minimal is the default.

Google describes minimal as matching a no-thinking path for most requests, although the model may still perform very limited reasoning on complex tasks.

This makes Flash-Lite a natural fit for classification, extraction, routing, translation, and other workloads where deep reasoning would add cost without materially improving the result.

When a subset of requests is more difficult, the application can increase the thinking level rather than moving every task to a larger model.

That gives developers two routing dimensions:

  1. choose the model tier;
  2. choose the thinking level inside that tier.

Thinking tokens are billed as output, so the correct setting should be chosen from evaluations rather than intuition.

Thinking levels
minimal · defaultlowmediumhigh

Routine high-volume task

Keep minimal for classification, extraction, translation, routing, and simple transformations when deeper reasoning does not improve acceptance.

Harder subtask

Raise thinking to low, medium, or high before escalating to a more expensive model when evaluation shows Flash-Lite can still complete the work.

08 / Tools

Flash-Lite Still Has a Broad Agent Tool Surface

Gemini 3.5 Flash-Lite is inexpensive, but it is not limited to plain text completions: Google lists search, maps, files, code execution, functions, structured outputs, URL context, caching, and Computer Use support.

Small Agents Can Still Act

The official model specification lists:

  • Google Search grounding;
  • Google Maps grounding;
  • File Search;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • context caching;
  • Computer Use in preview;
  • Batch API;
  • Flex inference;
  • Priority inference.

DeepMind also highlights tool use as a core capability and reports strong agentic-computer-use performance for the model class.

This enables low-cost agents that do more than classify text.

Examples include:

  • search-and-extract workflows;
  • lightweight browser automation;
  • file triage;
  • data enrichment;
  • validation agents;
  • high-volume internal tools;
  • narrow coding subagents.

Tool support does not remove the need for evaluation.

Track malformed arguments, wrong-tool selection, repeated calls, failed recovery, and whether the model stops when the task is complete.

Tool-capable Lite tier
  • Google Search

    Ground responses in current web information.

    Supported
  • Google Maps

    Use Maps grounding for location-aware workflows.

    Supported
  • File Search

    Retrieve relevant content from indexed file collections.

    Supported
  • Code execution

    Run code for calculations, transformations, and validation.

    Supported
  • Function calling

    Invoke application-defined tools and actions.

    Supported
  • Structured outputs

    Return machine-readable responses that follow an expected schema.

    Supported
  • Computer Use

    Preview support for graphical-interface interaction in supported environments.

    Supported
  • URL context

    Read and reason over content from supplied URLs.

    Supported

09 / Flash-Lite vs Flash

Gemini 3.5 Flash-Lite vs Gemini 3.5 Flash

Both models share a 1M-token multimodal context and broad Gemini tooling, but they optimize for different production constraints: Flash-Lite for throughput and cost, Flash for deeper coding and agentic capability.

Start With the Cheapest Model That Passes

Gemini 3.5 Flash-Lite costs:

  • $0.30 input;
  • $2.50 output.

Gemini 3.5 Flash costs:

  • $1.50 input;
  • $9.00 output.

That makes the standard Flash model five times more expensive on input and 3.6 times more expensive on output.

The capability tradeoff is equally important.

Gemini 3.5 Flash was launched for sustained frontier performance in coding, long-horizon agents, and difficult multi-step work.

Flash-Lite is optimized for high-throughput subagents, translation, parsing, extraction, and latency-sensitive reasoning.

Their default reasoning settings reflect that split:

  • Flash-Lite: minimal;
  • Flash: medium.

A useful routing strategy is therefore:

  1. attempt routine work on Flash-Lite;
  2. validate the result;
  3. increase Flash-Lite thinking for borderline cases;
  4. escalate difficult tasks to Flash only when needed.
Choose the tier
  1. 01

    Use Flash-Lite first

    Start with low-cost routing, extraction, classification, translation, parsing, and narrow subagent work.

  2. 02

    Tune reasoning

    Raise Flash-Lite from minimal when a task needs more reasoning but may not require a larger model.

  3. 03

    Validate quality

    Use deterministic checks, schema validation, tests, or task-specific acceptance criteria.

  4. 04

    Escalate selectively

    Route difficult coding, long-horizon planning, and persistent agent failures to Gemini 3.5 Flash or a newer model.

10 / 3.1 → 3.5 Lite

Gemini 3.5 Flash-Lite vs Gemini 3.1 Flash-Lite

Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite but substantially improves agentic coding, computer use, knowledge work, and long-context quality while remaining in the low-cost Lite tier.

The Upgrade Trades Slightly Higher Price for Much Stronger Execution

Google DeepMind’s July 2026 evaluations show large capability gains over Gemini 3.1 Flash-Lite across several agentic and coding benchmarks.

Examples include stronger results in:

  • diverse agentic coding;
  • terminal coding;
  • machine-learning engineering;
  • knowledge work;
  • computer use;
  • long-context retrieval.

The price also increases.

Gemini 3.1 Flash-Lite starts at $0.25 input and $1.50 output for standard text/image/video usage.

Gemini 3.5 Flash-Lite costs $0.30 input and $2.50 output.

That makes migration a price-performance decision rather than a pure cost reduction.

For workloads where 3.1 already meets acceptance criteria, its lower output price can remain attractive.

For tasks where 3.1 causes retries, failed agents, or weak computer-use behavior, 3.5 can reduce the total cost of a successful result despite the higher per-token rate.

Flash-Lite generation upgrade
  1. 01

    Replay failed 3.1 tasks

    Focus evaluation on agentic, coding, computer-use, and long-context cases where 3.1 Flash-Lite misses the acceptance threshold.

  2. 02

    Measure retry reduction

    Higher per-token pricing can still lower total cost if 3.5 avoids repeated attempts or stronger-model escalation.

  3. 03

    Keep routine traffic cheap

    If 3.1 already passes simple extraction or classification workloads, compare whether the 3.5 quality gain justifies migration.

  4. 04

    Compare completed-task economics

    Use quality, latency, retries, thinking tokens, and fallback cost together rather than comparing model price in isolation.

11 / Evaluation

Gemini 3.5 Flash-Lite Strengths and Limitations

Gemini 3.5 Flash-Lite is an efficiency specialist with unusually broad multimodal and agent capabilities for its price, but difficult reasoning and long-horizon work may still justify a larger Flash model.

Strengths

  • Very high throughput

    DeepMind describes it as the fastest 3.5-class model and reports roughly 350 output tokens per second on the Artificial Analysis Index.

  • Low standard cost

    $0.30 input and $2.50 output pricing makes it suitable for large classification, extraction, translation, routing, and subagent volumes.

  • Minimal thinking by default

    Latency-first reasoning behavior is the default, while low, medium, and high remain available for harder tasks.

  • Broad multimodal and tool support

    A 1M context plus text, image, video, audio, PDF, Search, Maps, files, code, functions, structured outputs, URLs, and Computer Use support sophisticated low-cost workflows.

What to consider

  • Not the deepest reasoning tier

    Flash-Lite is optimized for throughput and economics; difficult coding or long-horizon planning may benefit from Gemini Flash or Pro.

  • More expensive than 3.1 Flash-Lite

    The newer model improves capability but increases standard output pricing from $1.50 to $2.50 per million tokens.

  • Computer Use remains preview

    GUI automation requires permission boundaries, confirmation gates, monitoring, and regression testing.

  • Foundation-model limitations remain

    Google’s model card notes hallucination risk and occasional slowness or timeout behavior, so production validation is still required.

Scale the cheap path with evidence

Test Gemini 3.5 Flash-Lite on your high-volume workload

Replay extraction, classification, routing, translation, document parsing, subagent, and multimodal tasks to compare accepted-result rate, latency, thinking cost, tool-call reliability, and cost per completed task.

Start Free

Flash-Lite is most valuable when its lower latency and price survive real production evaluation. Measure retries, fallbacks, and validation overhead rather than selecting it from token price alone.

Common Questions

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google’s fastest and most cost-effective 3.5-class model, optimized for high-throughput subagents, document parsing, extraction, translation, multimodal processing, and latency-sensitive production workloads.

What is the Gemini 3.5 Flash-Lite model ID?

The stable Gemini API model ID is gemini-3.5-flash-lite.

When was Gemini 3.5 Flash-Lite released?

Google released Gemini 3.5 Flash-Lite as generally available on July 21, 2026.

Is Gemini 3.5 Flash-Lite stable?

Yes. Google lists gemini-3.5-flash-lite as Stable and generally available. Google Cloud lists retirement no earlier than July 21, 2027.

What is the Gemini 3.5 Flash-Lite context window?

Gemini 3.5 Flash-Lite supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.5 Flash-Lite support?

The model accepts text, images, video, audio, and PDF input and generates text output.

What is the Gemini 3.5 Flash-Lite knowledge cutoff?

Google DeepMind lists March 2026 as the knowledge cutoff, while noting that some domains may still reflect the broader Gemini 3 family’s January 2025 cutoff.

How much does Gemini 3.5 Flash-Lite cost?

Standard paid pricing is $0.30 per 1M input tokens and $2.50 per 1M output tokens, including thinking tokens.

How much does Gemini 3.5 Flash-Lite context caching cost?

Google lists standard cache reads at $0.03 per 1M cached tokens plus $1 per 1M cached tokens per hour of storage.

What are Gemini 3.5 Flash-Lite Batch API prices?

Batch pricing is $0.15 per 1M input tokens and $1.25 per 1M output tokens. Flex inference uses the same token rates.

Does Gemini 3.5 Flash-Lite have a free tier?

Yes. Google currently lists free standard token usage for Gemini 3.5 Flash-Lite in the Gemini Developer API, subject to provider limits and terms.

Does Gemini 3.5 Flash-Lite support thinking?

Yes. It supports minimal, low, medium, and high thinking levels.

What is the default thinking level of Gemini 3.5 Flash-Lite?

Minimal is the default thinking level, which prioritizes latency and cost for routine high-volume workloads.

How fast is Gemini 3.5 Flash-Lite?

Google DeepMind describes it as the fastest 3.5-class model and reports approximately 350 output tokens per second according to the Artificial Analysis Index.

Is Gemini 3.5 Flash-Lite good for document extraction?

Yes. Google explicitly positions the model for document parsing, simple data extraction, translation, classification, search, and other high-throughput processing workloads.

Can Gemini 3.5 Flash-Lite be used for agents?

Yes. Google positions it for high-volume agentic workflows and subagent tasks. It also supports function calling, code execution, Search, Maps, File Search, URL context, structured outputs, and Computer Use in preview.

Does Gemini 3.5 Flash-Lite support Computer Use?

Yes. Google’s current model documentation lists Computer Use as supported in preview.

What tools does Gemini 3.5 Flash-Lite support?

Google lists Search grounding, Maps grounding, File Search, code execution, function calling, structured outputs, URL context, context caching, and preview Computer Use. Batch, Flex, and Priority inference are also supported.

How is Gemini 3.5 Flash-Lite different from Gemini 3.5 Flash?

Both have a 1M input window and broad multimodal tool support, but Flash-Lite is optimized for low latency and cost with minimal thinking by default. Gemini 3.5 Flash is substantially more expensive and targets deeper coding, long-horizon agents, and frontier-level reasoning with medium thinking by default.

How is Gemini 3.5 Flash-Lite different from Gemini 3.1 Flash-Lite?

Gemini 3.5 Flash-Lite is based on 3.1 Flash-Lite but delivers much stronger agentic coding, computer use, knowledge work, and long-context performance in Google’s evaluations. Its standard token pricing is also higher.

Should I use Gemini 3.5 Flash-Lite for a new application?

It is a strong candidate when latency, request volume, or inference cost dominates the workload. Evaluate it first for routing, extraction, translation, document processing, multimodal batch jobs, and narrow subagents, then escalate only workloads that fail your quality threshold.

Model information

Last updated

Specifications, pricing, release status, thinking levels, throughput positioning, multimodal inputs, tool support, Computer Use, benchmark positioning, lifecycle, and migration guidance on this page are based on official Google Gemini API, Google DeepMind, and Google Cloud documentation.

Gemini 3.5 Flash-Lite — Pricing, 1M Context, High Throughput & Tools | EidoStack