Google Gemini model
Gemini 3.7 Flash
Google’s stable Flash workhorse for everyday coding, web development, multimodal knowledge work, agentic tool use, and reliable multi-step execution with a 1M-token input window.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $0.75
- intro · per 1M
- Cached input
- $0.075
- intro · per 1M
- Output
- $3.75
- intro · per 1M
01 / Overview
What Gemini 3.7 Flash Is
Gemini 3.7 Flash is Google’s stable Flash model for coding, agentic tool use, web development, multimodal knowledge work, and reliable multi-step execution.
The Flash Generation That Turned Into a Production Workhorse
Google released Gemini 3.7 Flash as generally available on August 13, 2026, only three weeks after Gemini 3.6 Flash.
At launch, Google described it as its most intelligent Flash workhorse yet for coding and agents.
The important distinction is not just benchmark intelligence. Gemini 3.7 Flash was positioned around practical developer workflows: debugging, issue resolution, first-pass code quality, web application generation, document comprehension, business automation, and tool-heavy agents.
The stable model ID is gemini-3.7-flash.
It supports up to 1,048,576 input tokens and 65,536 output tokens. Inputs can include text, images, video, audio, and PDFs, while output is text.
Gemini 3.7 Flash also supports low, medium, and high thinking levels, with medium as the default.
- Stable model ID:
gemini-3.7-flash. - GA since August 13, 2026.
- 1,048,576 input tokens.
- 65,536 maximum output tokens.
- Text, image, video, audio, and PDF input.
- Text output.
- Thinking levels: low, medium, high.
- Default thinking: medium.
- Provider
- Family
- Gemini 3.7 Flash
- Model ID
- gemini-3.7-flash
- Status
- Stable · GA
- Knowledge cutoff
- Jan 2025
- Input
- Text, Image, Video, Audio, PDF
- Output
- Text
02 / Production coding
Production Coding With Better First-Pass Accuracy
Gemini 3.7 Flash was launched around stronger software engineering: better debugging, issue resolution, production-ready code, and more disciplined execution than the preceding Flash generation.
Fewer Retries Can Matter More Than Fewer Tokens
Coding cost is not only the cost of one generated answer.
A development agent may inspect repository files, reason about dependencies, produce a patch, run tests, analyze failures, revise its plan, and repeat until the change passes validation.
Google highlighted improved first-pass code accuracy in Gemini 3.7 Flash and stronger performance on production-code and software-engineering evaluations.
That changes the relevant production metric.
A model that uses slightly more reasoning but produces a correct implementation in one or two attempts can be cheaper than a more token-efficient model that repeatedly fails tests.
Useful metrics include:
- first-pass acceptance rate;
- tests passed after the first patch;
- number of corrective prompts;
- tool-call count;
- repeated file reads;
- failed execution loops;
- latency to accepted code;
- total tokens per completed task.
Gemini 3.7 Flash is therefore especially useful as a production coding baseline against which newer Flash models can be measured.
- 01
Understand
Read code, requirements, errors, and surrounding repository context before changing files.
- 02
Implement
Generate production-oriented changes rather than isolated code snippets.
- 03
Validate
Use tools, code execution, and test results to determine whether the implementation actually works.
- 04
Repair
Adapt to roadblocks, revise failed approaches, and reduce manual developer intervention.
03 / Web development
Web Development Is a Distinct Gemini 3.7 Flash Strength
Google specifically highlighted web development as a major improvement in Gemini 3.7 Flash, including more functional layouts, better design adherence, and feature-complete applications generated in fewer prompts.
Evaluate the Whole Frontend, Not Only the JSX
Web-development quality spans several dimensions at once.
A generated page can compile and still be unusable. It can look visually close to a reference while missing interactions. It can implement features correctly but ignore the design system.
Gemini 3.7 Flash is particularly relevant for workflows that combine:
- screenshot-to-code;
- image-reference implementation;
- design-system adherence;
- responsive layout generation;
- landing pages;
- interactive components;
- frontend feature implementation;
- iterative browser or Computer Use validation.
Because Gemini accepts images directly, a developer can supply screenshots or visual references alongside code and written requirements.
For production evaluation, measure visual parity, functional completeness, responsiveness, accessibility, build success, and the number of correction prompts required before acceptance.
This web-development emphasis gives Gemini 3.7 Flash a different SEO and evaluation identity from Gemini 3.8 Flash, whose strongest positioning moves further toward long-horizon software engineering and autonomous agent execution.
- 01
Reference understanding
Interpret screenshots, images, layouts, and design-system examples alongside written requirements.
- 02
Functional implementation
Generate interfaces that work rather than merely resemble a visual reference.
- 03
Design adherence
Preserve layout, hierarchy, spacing, and visual intent across generated components.
- 04
Iteration efficiency
Track how many prompts and browser-validation loops are needed before the UI is accepted.
04 / Knowledge work
Multimodal Knowledge Work Beyond Coding
Gemini 3.7 Flash also improved document comprehension, knowledge-dense reasoning, and real-world business workflow automation.
A Flash Model for Finance, Legal, Scientific, and Enterprise Material
Google highlighted stronger performance in knowledge-heavy domains such as finance, law, and biosciences.
That matters because enterprise tasks frequently combine several kinds of information:
- long PDFs;
- charts and images;
- structured files;
- policy or legal text;
- live web context;
- internal tools;
- business actions.
Gemini 3.7 Flash can keep these inputs inside a large multimodal context and combine them with built-in retrieval, search, URL context, code execution, and function calling.
A useful evaluation should separate document comprehension from workflow execution.
The model may correctly understand an annual report but still select the wrong downstream action. Or it may call the right tool but extract the wrong value from a table.
Measure factual extraction, cross-document consistency, tool correctness, citation grounding, structured-output validity, and whether the complete workflow finishes without manual repair.
- 01
PDF comprehension
Analyze long reports, policies, scientific material, and other document-heavy inputs.
- 02
Knowledge synthesis
Combine evidence from text, charts, images, URLs, and retrieved sources.
- 03
Business automation
Turn analysis into tool calls, updates, structured outputs, or downstream actions.
- 04
Multimodal review
Evaluate text and visual information together rather than splitting the workflow across separate models.
05 / Pricing
Gemini 3.7 Flash Pricing
Gemini 3.7 Flash uses the same introductory and scheduled standard pricing as Gemini 3.8 Flash: $0.75/$3.75 through the end of 2026, then $1.50/$7.50 beginning January 1, 2027.
Model 2027 Economics Before You Commit
Through December 31, 2026, Google lists:
- $0.75 per 1M input tokens.
- $3.75 per 1M output tokens, including thinking tokens.
- $0.075 per 1M cached-context tokens.
- $0.50 per 1M cached tokens per hour of storage.
Starting January 1, 2027, Google lists:
- $1.50 per 1M input tokens.
- $7.50 per 1M output tokens.
- $0.15 per 1M cached-context tokens.
- $1.00 per 1M cached tokens per hour of storage.
Batch and Flex inference use discounted rates, while Priority inference provides a premium serving option.
This scheduled pricing change is important for model evaluation in late 2026.
A project that launches under the introductory rate but runs throughout 2027 should calculate unit economics using the future standard rate as well.
1M tokens · USD
- Input
- $0.75
- Cached input
- $0.075
- Output + thinking
- $3.75
Example: 20K input + 4K output
- Input cost
- $0.0150
- Output cost
- $0.0150
- Estimated introductory total
- $0.0300
06 / Multimodal context
1,048,576 Tokens of Native Multimodal Input
Gemini 3.7 Flash accepts up to 1,048,576 input tokens across text, images, video, audio, and PDFs, with a 65,536-token output ceiling.
Long Context Is a Working Set, Not Just a Document Limit
The context window can contain much more than conversation text.
A coding workflow might combine repository files, screenshots, architecture PDFs, logs, user requirements, tool results, and previous agent turns.
A knowledge-work task might combine a long financial report, charts, recorded audio, supplemental URLs, and structured instructions.
This native multimodality reduces the need to split each media type into a separate model pipeline.
However, a large context window can become expensive and noisy when the application sends irrelevant state.
Use retrieval, caching, selective history, and explicit context budgeting to preserve the information that actually contributes to the task.
Input context
1,048,576
Max output
65,536
Text, images, video, audio, PDFs, instructions, history, and tool results all consume the shared input context.
07 / Thinking
Low, Medium, and High Thinking
Gemini 3.7 Flash uses dynamic thinking by default and supports three explicit reasoning levels: low, medium, and high.
Medium Is the Default Production Balance
Google documents medium as the default for Gemini 3.7 Flash.
Use low when time-to-first-answer and token cost matter more than deep deliberation.
Use medium for the general case: coding, agentic tool use, document work, and multi-step execution.
Use high when the workload consistently benefits from deeper reasoning or more careful tool orchestration.
minimal is not supported and returns an API validation error.
Thinking tokens are included in output billing, so the setting affects both quality and cost.
Gemini 3.7 Flash was designed to think more diligently than 3.6 Flash on difficult workflows. Google described this as more effort in multi-step planning and tool calling, with the goal of reducing manual oversight and retries.
Fast interaction
Use low when latency matters and your evals show that additional reasoning does not improve the accepted result.
Complex coding or tools
Use medium or high when better planning, debugging, and tool orchestration reduce retries and manual intervention.
08 / Tools
Search, Code, Files, Functions, and Computer Use
Gemini 3.7 Flash exposes the same broad Gemini 3 application surface that lets the model retrieve current information, execute code, search files, invoke functions, and interact with graphical interfaces.
Tool Support Is Part of the Model’s Production Identity
The official model specification lists support for Google Search grounding, Google Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and Computer Use in preview.
This allows a single agent to move between reasoning and action.
A coding workflow can inspect files, call an application tool, execute code, review the result, and produce structured output.
A knowledge workflow can read a PDF, ground current information through Search, inspect a URL, and call a business operation.
Computer Use extends the same pattern to browser and GUI workflows.
Tool support should still be evaluated at the trajectory level. Measure wrong-tool calls, invalid arguments, unnecessary calls, recovery from failures, and the number of steps required to complete the task.
- Supported
Google Search
Ground answers in current web information and return source-backed results.
- Supported
Google Maps
Ground location-aware tasks in Maps data.
- Supported
File Search
Retrieve relevant information from indexed file collections.
- Supported
Code execution
Run code during reasoning and validation workflows.
- Supported
Function calling
Invoke application-defined tools and business operations.
- Supported
Structured outputs
Return machine-readable responses that follow a requested schema.
- Supported
Computer Use
Preview support for agents that interact with graphical user interfaces.
- Supported
URL context
Read and reason over content from supplied URLs.
09 / 3.7 → 3.8
Gemini 3.7 Flash vs Gemini 3.8 Flash
Gemini 3.8 Flash keeps the same 1M/64K capacity, multimodal inputs, thinking levels, and 2026 token prices, but moves the Flash tier further toward long-horizon software engineering and autonomous agent execution.
The Upgrade Is About Completion Reliability More Than Raw Capacity
The obvious specifications are almost unchanged.
Both models provide:
- 1,048,576 input tokens;
- 65,536 output tokens;
- low, medium, and high thinking;
- multimodal input;
- broad built-in tool support;
- the same introductory and scheduled standard token pricing.
The key difference is behavioral.
Google positions 3.7 Flash around everyday coding, agentic tool use, web development, first-pass code accuracy, knowledge work, and reliable multi-step execution.
Gemini 3.8 Flash moves the same Flash economics further toward long-horizon software engineering, autonomous agents, stronger multi-step problem solving, and more deliberate verification.
Google also notes that 3.8 can use more tokens on difficult tasks because it may take smaller reasoning steps, invoke tools iteratively, and verify its work more extensively.
That makes 3.7 a valuable comparison baseline.
If 3.7 already passes your coding or web-development acceptance threshold, 3.8’s additional token usage may not be necessary. If long tasks frequently fail or require human rescue, 3.8 may reduce total task cost despite consuming more tokens.
- 01
Keep the same capacity expectations
Both models use a 1,048,576-token input limit and 65,536-token output limit.
- 02
Replay long trajectories
Focus migration evals on tasks where 3.7 fails late, loops, loses state, or requires human intervention.
- 03
Measure token expansion
3.8 may deliberately spend more tokens on difficult work, so compare cost per successful task instead of equal request length.
- 04
Keep 3.7 where it already passes
There is little value in paying for extra reasoning if 3.7 reliably clears the product’s acceptance threshold.
10 / Evaluation
Gemini 3.7 Flash Strengths and Limitations
Gemini 3.7 Flash remains a strong stable production baseline for coding, web development, multimodal knowledge work, and agents, but it is no longer Google’s most capable Flash model for long-horizon execution.
Strengths
Production coding focus
Google launched 3.7 Flash around stronger debugging, issue resolution, first-pass accuracy, and production-ready code.
Distinct web-development strength
UI generation, reference-image adherence, functional layouts, and feature-complete web apps are central to the model’s launch positioning.
Broad multimodal knowledge work
A 1M-token context can combine PDFs, charts, images, audio, video, code, URLs, and tool results.
Mature tool ecosystem
Search, Maps, File Search, code execution, function calling, structured output, URL context, caching, and Computer Use support real agent workflows.
What to consider
3.8 is stronger for long-horizon work
Google positions Gemini 3.8 Flash as the newer choice for long-running software engineering, autonomous agents, and harder multi-step execution.
Introductory pricing expires
The current $0.75/$3.75 standard rate is scheduled to become $1.50/$7.50 on January 1, 2027.
No minimal thinking
Low is the least intensive supported reasoning level; minimal returns an error.
Computer Use remains preview
GUI-agent workflows need additional safeguards, regression tests, and production monitoring.
Benchmark the production baseline
Test Gemini 3.7 Flash on your real coding and agent workflows
Replay code generation, debugging, UI implementation, PDF analysis, tool-use, and multimodal tasks to compare first-pass accuracy, retries, token usage, latency, and cost against Gemini 3.8 Flash and other models.
Start FreeGemini 3.7 Flash remains a useful stable baseline when you want to measure whether 3.8’s stronger long-horizon behavior actually improves your workload.
Common Questions
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is Google’s stable Flash model for everyday coding, agentic tool use, web development, multimodal knowledge work, and reliable multi-step execution.
What is the Gemini 3.7 Flash model ID?
The stable Gemini API model ID is gemini-3.7-flash.
When was Gemini 3.7 Flash released?
Google released Gemini 3.7 Flash as generally available on August 13, 2026.
Is Gemini 3.7 Flash deprecated?
No. Google currently lists gemini-3.7-flash as a stable model and has not announced a shutdown date.
What is the Gemini 3.7 Flash context window?
Gemini 3.7 Flash supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3.7 Flash support?
The model accepts text, images, video, audio, and PDF input and generates text output.
What is the Gemini 3.7 Flash knowledge cutoff?
Google documents January 2025 as the Gemini 3 knowledge cutoff. Gemini 3.7 Flash can use Google Search grounding when current information is required.
How much does Gemini 3.7 Flash cost?
Through December 31, 2026, standard paid pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, including thinking tokens. Starting January 1, 2027, Google lists $1.50 input and $7.50 output per million tokens.
How much does Gemini 3.7 Flash context caching cost?
Through December 31, 2026, cached context costs $0.075 per 1M tokens plus $0.50 per 1M tokens per hour of storage. Starting January 1, 2027, Google lists $0.15 per 1M cached tokens plus $1.00 per 1M tokens per hour of storage.
Does Gemini 3.7 Flash support thinking?
Yes. Gemini 3.7 Flash supports low, medium, and high thinking levels. Medium is the default.
Does Gemini 3.7 Flash support minimal thinking?
No. Google states that minimal is unsupported and returns an API validation error. Use low for the least intensive supported thinking level.
Is Gemini 3.7 Flash good for coding?
Coding is one of the model’s primary use cases. Google highlighted improvements in debugging, issue resolution, first-pass code accuracy, production-ready code, and developer instruction following.
Is Gemini 3.7 Flash good for web development?
Yes. Google specifically highlighted web development, including functional layouts, feature-complete apps, UI generation, and stronger adherence to screenshots, images, and design-system references.
Does Gemini 3.7 Flash support Computer Use?
Yes, in preview. The official model specification lists Computer Use as supported.
What tools does Gemini 3.7 Flash support?
The official model specification lists Google Search grounding, Google Maps grounding, File Search, code execution, URL context, function calling, structured outputs, context caching, and preview Computer Use.
How is Gemini 3.7 Flash different from Gemini 3.8 Flash?
Both have the same 1M input window, 65,536 output limit, thinking levels, multimodal inputs, and 2026 pricing. Gemini 3.7 Flash is strongly positioned around everyday coding, web development, and knowledge work, while Google positions 3.8 Flash for stronger long-horizon software engineering, autonomous agents, and more demanding multi-step execution.
Should I use Gemini 3.7 Flash or Gemini 3.8 Flash?
Use workload evaluations. Keep Gemini 3.7 Flash where it already meets quality, latency, and cost targets. Evaluate Gemini 3.8 Flash for tasks where long trajectories fail, agents loop, or difficult engineering work requires more reliable planning and verification.
Model information
Last updated
Specifications, release status, pricing, thinking levels, multimodal inputs, coding and web-development positioning, built-in tools, Computer Use, and migration guidance on this page are based on official Google Gemini API and Google model documentation.