Google Gemini model

Gemini 3.7 Flash

Google’s stable Flash workhorse for everyday coding, web development, multimodal knowledge work, agentic tool use, and reliable multi-step execution with a 1M-token input window.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$0.75
intro · per 1M
Cached input
$0.075
intro · per 1M
Output
$3.75
intro · per 1M

01 / Overview

What Gemini 3.7 Flash Is

Gemini 3.7 Flash is Google’s stable Flash model for coding, agentic tool use, web development, multimodal knowledge work, and reliable multi-step execution.

The Flash Generation That Turned Into a Production Workhorse

Google released Gemini 3.7 Flash as generally available on August 13, 2026, only three weeks after Gemini 3.6 Flash.

At launch, Google described it as its most intelligent Flash workhorse yet for coding and agents.

The important distinction is not just benchmark intelligence. Gemini 3.7 Flash was positioned around practical developer workflows: debugging, issue resolution, first-pass code quality, web application generation, document comprehension, business automation, and tool-heavy agents.

The stable model ID is gemini-3.7-flash.

It supports up to 1,048,576 input tokens and 65,536 output tokens. Inputs can include text, images, video, audio, and PDFs, while output is text.

Gemini 3.7 Flash also supports low, medium, and high thinking levels, with medium as the default.

  • Stable model ID: gemini-3.7-flash.
  • GA since August 13, 2026.
  • 1,048,576 input tokens.
  • 65,536 maximum output tokens.
  • Text, image, video, audio, and PDF input.
  • Text output.
  • Thinking levels: low, medium, high.
  • Default thinking: medium.
Model profile
Provider
Google
Family
Gemini 3.7 Flash
Model ID
gemini-3.7-flash
Status
Stable · GA
Knowledge cutoff
Jan 2025
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / Production coding

Production Coding With Better First-Pass Accuracy

Gemini 3.7 Flash was launched around stronger software engineering: better debugging, issue resolution, production-ready code, and more disciplined execution than the preceding Flash generation.

Fewer Retries Can Matter More Than Fewer Tokens

Coding cost is not only the cost of one generated answer.

A development agent may inspect repository files, reason about dependencies, produce a patch, run tests, analyze failures, revise its plan, and repeat until the change passes validation.

Google highlighted improved first-pass code accuracy in Gemini 3.7 Flash and stronger performance on production-code and software-engineering evaluations.

That changes the relevant production metric.

A model that uses slightly more reasoning but produces a correct implementation in one or two attempts can be cheaper than a more token-efficient model that repeatedly fails tests.

Useful metrics include:

  • first-pass acceptance rate;
  • tests passed after the first patch;
  • number of corrective prompts;
  • tool-call count;
  • repeated file reads;
  • failed execution loops;
  • latency to accepted code;
  • total tokens per completed task.

Gemini 3.7 Flash is therefore especially useful as a production coding baseline against which newer Flash models can be measured.

Coding workflow
  1. 01

    Understand

    Read code, requirements, errors, and surrounding repository context before changing files.

  2. 02

    Implement

    Generate production-oriented changes rather than isolated code snippets.

  3. 03

    Validate

    Use tools, code execution, and test results to determine whether the implementation actually works.

  4. 04

    Repair

    Adapt to roadblocks, revise failed approaches, and reduce manual developer intervention.

03 / Web development

Web Development Is a Distinct Gemini 3.7 Flash Strength

Google specifically highlighted web development as a major improvement in Gemini 3.7 Flash, including more functional layouts, better design adherence, and feature-complete applications generated in fewer prompts.

Evaluate the Whole Frontend, Not Only the JSX

Web-development quality spans several dimensions at once.

A generated page can compile and still be unusable. It can look visually close to a reference while missing interactions. It can implement features correctly but ignore the design system.

Gemini 3.7 Flash is particularly relevant for workflows that combine:

  • screenshot-to-code;
  • image-reference implementation;
  • design-system adherence;
  • responsive layout generation;
  • landing pages;
  • interactive components;
  • frontend feature implementation;
  • iterative browser or Computer Use validation.

Because Gemini accepts images directly, a developer can supply screenshots or visual references alongside code and written requirements.

For production evaluation, measure visual parity, functional completeness, responsiveness, accessibility, build success, and the number of correction prompts required before acceptance.

This web-development emphasis gives Gemini 3.7 Flash a different SEO and evaluation identity from Gemini 3.8 Flash, whose strongest positioning moves further toward long-horizon software engineering and autonomous agent execution.

UI and app generation
  1. 01

    Reference understanding

    Interpret screenshots, images, layouts, and design-system examples alongside written requirements.

  2. 02

    Functional implementation

    Generate interfaces that work rather than merely resemble a visual reference.

  3. 03

    Design adherence

    Preserve layout, hierarchy, spacing, and visual intent across generated components.

  4. 04

    Iteration efficiency

    Track how many prompts and browser-validation loops are needed before the UI is accepted.

04 / Knowledge work

Multimodal Knowledge Work Beyond Coding

Gemini 3.7 Flash also improved document comprehension, knowledge-dense reasoning, and real-world business workflow automation.

A Flash Model for Finance, Legal, Scientific, and Enterprise Material

Google highlighted stronger performance in knowledge-heavy domains such as finance, law, and biosciences.

That matters because enterprise tasks frequently combine several kinds of information:

  • long PDFs;
  • charts and images;
  • structured files;
  • policy or legal text;
  • live web context;
  • internal tools;
  • business actions.

Gemini 3.7 Flash can keep these inputs inside a large multimodal context and combine them with built-in retrieval, search, URL context, code execution, and function calling.

A useful evaluation should separate document comprehension from workflow execution.

The model may correctly understand an annual report but still select the wrong downstream action. Or it may call the right tool but extract the wrong value from a table.

Measure factual extraction, cross-document consistency, tool correctness, citation grounding, structured-output validity, and whether the complete workflow finishes without manual repair.

Documents and enterprise workflows
  1. 01

    PDF comprehension

    Analyze long reports, policies, scientific material, and other document-heavy inputs.

  2. 02

    Knowledge synthesis

    Combine evidence from text, charts, images, URLs, and retrieved sources.

  3. 03

    Business automation

    Turn analysis into tool calls, updates, structured outputs, or downstream actions.

  4. 04

    Multimodal review

    Evaluate text and visual information together rather than splitting the workflow across separate models.

05 / Pricing

Gemini 3.7 Flash Pricing

Gemini 3.7 Flash uses the same introductory and scheduled standard pricing as Gemini 3.8 Flash: $0.75/$3.75 through the end of 2026, then $1.50/$7.50 beginning January 1, 2027.

Model 2027 Economics Before You Commit

Through December 31, 2026, Google lists:

  • $0.75 per 1M input tokens.
  • $3.75 per 1M output tokens, including thinking tokens.
  • $0.075 per 1M cached-context tokens.
  • $0.50 per 1M cached tokens per hour of storage.

Starting January 1, 2027, Google lists:

  • $1.50 per 1M input tokens.
  • $7.50 per 1M output tokens.
  • $0.15 per 1M cached-context tokens.
  • $1.00 per 1M cached tokens per hour of storage.

Batch and Flex inference use discounted rates, while Priority inference provides a premium serving option.

This scheduled pricing change is important for model evaluation in late 2026.

A project that launches under the introductory rate but runs throughout 2027 should calculate unit economics using the future standard rate as well.

Introductory and standard pricing

1M tokens · USD

Input
$0.75
Cached input
$0.075
Output + thinking
$3.75

Example: 20K input + 4K output

Input cost
$0.0150
Output cost
$0.0150
Estimated introductory total
$0.0300

06 / Multimodal context

1,048,576 Tokens of Native Multimodal Input

Gemini 3.7 Flash accepts up to 1,048,576 input tokens across text, images, video, audio, and PDFs, with a 65,536-token output ceiling.

Long Context Is a Working Set, Not Just a Document Limit

The context window can contain much more than conversation text.

A coding workflow might combine repository files, screenshots, architecture PDFs, logs, user requirements, tool results, and previous agent turns.

A knowledge-work task might combine a long financial report, charts, recorded audio, supplemental URLs, and structured instructions.

This native multimodality reduces the need to split each media type into a separate model pipeline.

However, a large context window can become expensive and noisy when the application sends irrelevant state.

Use retrieval, caching, selective history, and explicit context budgeting to preserve the information that actually contributes to the task.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Text, images, video, audio, PDFs, instructions, history, and tool results all consume the shared input context.

07 / Thinking

Low, Medium, and High Thinking

Gemini 3.7 Flash uses dynamic thinking by default and supports three explicit reasoning levels: low, medium, and high.

Medium Is the Default Production Balance

Google documents medium as the default for Gemini 3.7 Flash.

Use low when time-to-first-answer and token cost matter more than deep deliberation.

Use medium for the general case: coding, agentic tool use, document work, and multi-step execution.

Use high when the workload consistently benefits from deeper reasoning or more careful tool orchestration.

minimal is not supported and returns an API validation error.

Thinking tokens are included in output billing, so the setting affects both quality and cost.

Gemini 3.7 Flash was designed to think more diligently than 3.6 Flash on difficult workflows. Google described this as more effort in multi-step planning and tool calling, with the goal of reducing manual oversight and retries.

Thinking levels
lowmedium · defaulthigh

Fast interaction

Use low when latency matters and your evals show that additional reasoning does not improve the accepted result.

Complex coding or tools

Use medium or high when better planning, debugging, and tool orchestration reduce retries and manual intervention.

08 / Tools

Search, Code, Files, Functions, and Computer Use

Gemini 3.7 Flash exposes the same broad Gemini 3 application surface that lets the model retrieve current information, execute code, search files, invoke functions, and interact with graphical interfaces.

Tool Support Is Part of the Model’s Production Identity

The official model specification lists support for Google Search grounding, Google Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and Computer Use in preview.

This allows a single agent to move between reasoning and action.

A coding workflow can inspect files, call an application tool, execute code, review the result, and produce structured output.

A knowledge workflow can read a PDF, ground current information through Search, inspect a URL, and call a business operation.

Computer Use extends the same pattern to browser and GUI workflows.

Tool support should still be evaluated at the trajectory level. Measure wrong-tool calls, invalid arguments, unnecessary calls, recovery from failures, and the number of steps required to complete the task.

Built-in capabilities
  • Google Search

    Ground answers in current web information and return source-backed results.

    Supported
  • Google Maps

    Ground location-aware tasks in Maps data.

    Supported
  • File Search

    Retrieve relevant information from indexed file collections.

    Supported
  • Code execution

    Run code during reasoning and validation workflows.

    Supported
  • Function calling

    Invoke application-defined tools and business operations.

    Supported
  • Structured outputs

    Return machine-readable responses that follow a requested schema.

    Supported
  • Computer Use

    Preview support for agents that interact with graphical user interfaces.

    Supported
  • URL context

    Read and reason over content from supplied URLs.

    Supported

09 / 3.7 → 3.8

Gemini 3.7 Flash vs Gemini 3.8 Flash

Gemini 3.8 Flash keeps the same 1M/64K capacity, multimodal inputs, thinking levels, and 2026 token prices, but moves the Flash tier further toward long-horizon software engineering and autonomous agent execution.

The Upgrade Is About Completion Reliability More Than Raw Capacity

The obvious specifications are almost unchanged.

Both models provide:

  • 1,048,576 input tokens;
  • 65,536 output tokens;
  • low, medium, and high thinking;
  • multimodal input;
  • broad built-in tool support;
  • the same introductory and scheduled standard token pricing.

The key difference is behavioral.

Google positions 3.7 Flash around everyday coding, agentic tool use, web development, first-pass code accuracy, knowledge work, and reliable multi-step execution.

Gemini 3.8 Flash moves the same Flash economics further toward long-horizon software engineering, autonomous agents, stronger multi-step problem solving, and more deliberate verification.

Google also notes that 3.8 can use more tokens on difficult tasks because it may take smaller reasoning steps, invoke tools iteratively, and verify its work more extensively.

That makes 3.7 a valuable comparison baseline.

If 3.7 already passes your coding or web-development acceptance threshold, 3.8’s additional token usage may not be necessary. If long tasks frequently fail or require human rescue, 3.8 may reduce total task cost despite consuming more tokens.

Migration comparison
  1. 01

    Keep the same capacity expectations

    Both models use a 1,048,576-token input limit and 65,536-token output limit.

  2. 02

    Replay long trajectories

    Focus migration evals on tasks where 3.7 fails late, loops, loses state, or requires human intervention.

  3. 03

    Measure token expansion

    3.8 may deliberately spend more tokens on difficult work, so compare cost per successful task instead of equal request length.

  4. 04

    Keep 3.7 where it already passes

    There is little value in paying for extra reasoning if 3.7 reliably clears the product’s acceptance threshold.

10 / Evaluation

Gemini 3.7 Flash Strengths and Limitations

Gemini 3.7 Flash remains a strong stable production baseline for coding, web development, multimodal knowledge work, and agents, but it is no longer Google’s most capable Flash model for long-horizon execution.

Strengths

  • Production coding focus

    Google launched 3.7 Flash around stronger debugging, issue resolution, first-pass accuracy, and production-ready code.

  • Distinct web-development strength

    UI generation, reference-image adherence, functional layouts, and feature-complete web apps are central to the model’s launch positioning.

  • Broad multimodal knowledge work

    A 1M-token context can combine PDFs, charts, images, audio, video, code, URLs, and tool results.

  • Mature tool ecosystem

    Search, Maps, File Search, code execution, function calling, structured output, URL context, caching, and Computer Use support real agent workflows.

What to consider

  • 3.8 is stronger for long-horizon work

    Google positions Gemini 3.8 Flash as the newer choice for long-running software engineering, autonomous agents, and harder multi-step execution.

  • Introductory pricing expires

    The current $0.75/$3.75 standard rate is scheduled to become $1.50/$7.50 on January 1, 2027.

  • No minimal thinking

    Low is the least intensive supported reasoning level; minimal returns an error.

  • Computer Use remains preview

    GUI-agent workflows need additional safeguards, regression tests, and production monitoring.

Benchmark the production baseline

Test Gemini 3.7 Flash on your real coding and agent workflows

Replay code generation, debugging, UI implementation, PDF analysis, tool-use, and multimodal tasks to compare first-pass accuracy, retries, token usage, latency, and cost against Gemini 3.8 Flash and other models.

Start Free

Gemini 3.7 Flash remains a useful stable baseline when you want to measure whether 3.8’s stronger long-horizon behavior actually improves your workload.

Common Questions

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is Google’s stable Flash model for everyday coding, agentic tool use, web development, multimodal knowledge work, and reliable multi-step execution.

What is the Gemini 3.7 Flash model ID?

The stable Gemini API model ID is gemini-3.7-flash.

When was Gemini 3.7 Flash released?

Google released Gemini 3.7 Flash as generally available on August 13, 2026.

Is Gemini 3.7 Flash deprecated?

No. Google currently lists gemini-3.7-flash as a stable model and has not announced a shutdown date.

What is the Gemini 3.7 Flash context window?

Gemini 3.7 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.7 Flash support?

The model accepts text, images, video, audio, and PDF input and generates text output.

What is the Gemini 3.7 Flash knowledge cutoff?

Google documents January 2025 as the Gemini 3 knowledge cutoff. Gemini 3.7 Flash can use Google Search grounding when current information is required.

How much does Gemini 3.7 Flash cost?

Through December 31, 2026, standard paid pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, Google lists $1.50 input and $7.50 output per million tokens.

How much does Gemini 3.7 Flash context caching cost?

Through December 31, 2026, cached context costs $0.075 per 1M tokens plus $0.50 per 1M tokens per hour of storage. Starting January 1, 2027, Google lists $0.15 per 1M cached tokens plus $1.00 per 1M tokens per hour of storage.

Does Gemini 3.7 Flash support thinking?

Yes. Gemini 3.7 Flash supports low, medium, and high thinking levels. Medium is the default.

Does Gemini 3.7 Flash support minimal thinking?

No. Google states that minimal is unsupported and returns an API validation error. Use low for the least intensive supported thinking level.

Is Gemini 3.7 Flash good for coding?

Coding is one of the model’s primary use cases. Google highlighted improvements in debugging, issue resolution, first-pass code accuracy, production-ready code, and developer instruction following.

Is Gemini 3.7 Flash good for web development?

Yes. Google specifically highlighted web development, including functional layouts, feature-complete apps, UI generation, and stronger adherence to screenshots, images, and design-system references.

Does Gemini 3.7 Flash support Computer Use?

Yes, in preview. The official model specification lists Computer Use as supported.

What tools does Gemini 3.7 Flash support?

The official model specification lists Google Search grounding, Google Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use.

How is Gemini 3.7 Flash different from Gemini 3.8 Flash?

Both have the same 1M input window, 65,536 output limit, thinking levels, multimodal inputs, and 2026 pricing. Gemini 3.7 Flash is strongly positioned around everyday coding, web development, and knowledge work, while Google positions 3.8 Flash for stronger long-horizon software engineering, autonomous agents, and more demanding multi-step execution.

Should I use Gemini 3.7 Flash or Gemini 3.8 Flash?

Use workload evaluations. Keep Gemini 3.7 Flash where it already meets quality, latency, and cost targets. Evaluate Gemini 3.8 Flash for tasks where long trajectories fail, agents loop, or difficult engineering work requires more reliable planning and verification.

Model information

Last updated

Specifications, release status, pricing, thinking levels, multimodal inputs, coding and web-development positioning, built-in tools, Computer Use, and migration guidance on this page are based on official Google Gemini API and Google model documentation.

Gemini 3.7 Flash — Pricing, 1M Context, Coding, Web Development & Tools | EidoStack