Google Gemini preview model

Gemini 3.1 Pro Preview

Google’s preview Pro-tier Gemini model for complex multimodal reasoning, software engineering, vibe-coding, and agentic workflows that need precise tool use and reliable multi-step execution.

Input context
1.05M
tokens
Max output
65.5K
tokens
Input
$2.00
≤200K prompt · per 1M
Cached input
$0.20
≤200K prompt · per 1M
Output
$12.00
≤200K prompt · per 1M

01 / Overview

What Gemini 3.1 Pro Preview Is

Gemini 3.1 Pro Preview is Google’s preview Pro-tier model for complex tasks that need broad world knowledge, advanced multimodal reasoning, software-engineering capability, and reliable agentic execution.

A Pro Model Focused on Reliability, Not Throughput

Google describes Gemini 3.1 Pro Preview as a refinement of the Gemini 3 Pro series with better thinking, improved token efficiency, and a more grounded, factually consistent experience.

Its product role is therefore different from Gemini Flash.

Flash models optimize the capability-throughput-cost balance. Gemini 3.1 Pro Preview is the model to evaluate when the task is difficult enough that reasoning quality, factual consistency, software-engineering behavior, and precise tool usage matter more than lowest token price.

Google released gemini-3.1-pro-preview on February 19, 2026.

The model accepts text, images, video, audio, and PDFs and returns text. It supports up to 1,048,576 input tokens and up to 65,536 output tokens.

It also supports dynamic thinking, Google Search grounding, Google Maps grounding, code execution, function calling, structured outputs, URL context, context caching, and several inference tiers.

  • Model ID: gemini-3.1-pro-preview.
  • Launch stage: Public preview.
  • Released February 19, 2026.
  • 1,048,576 input tokens.
  • 65,536 maximum output tokens.
  • Text, image, video, audio, and PDF input.
  • Text output.
  • Default thinking: high.
  • Separate custom-tools endpoint available.
Model profile
Provider
Google
Family
Gemini 3.1 Pro
Model ID
gemini-3.1-pro-preview
Launch stage
Public preview
Knowledge cutoff
Jan 2025
Input
Text, Image, Video, Audio, PDF
Output
Text

02 / Reasoning & reliability

Better Thinking and More Grounded Multi-Step Execution

Google positions Gemini 3.1 Pro Preview around better thinking, improved token efficiency, stronger grounding, factual consistency, and more reliable execution across complex real-world tasks.

Evaluate the Entire Reasoning Trajectory

A difficult production task is rarely one prompt followed by one answer.

The model may need to understand a large context, form a plan, search for external evidence, call tools, evaluate results, revise its approach, and produce a final answer that remains consistent with everything it observed.

Failures can happen at any point:

  • the initial plan can be wrong;
  • a tool can be selected incorrectly;
  • valid evidence can be ignored;
  • a later answer can contradict an earlier observation;
  • the model can spend too many tokens before reaching a useful result.

Gemini 3.1 Pro Preview is specifically intended to improve this class of behavior.

For evaluation, measure:

  • final-answer correctness;
  • factual consistency across long responses;
  • successful tool selection;
  • valid tool arguments;
  • recovery after failed calls;
  • number of unnecessary steps;
  • thinking-token usage;
  • time to accepted result.

The relevant unit is not “one completion.” It is the complete reasoning-and-action trajectory.

Pro-tier evaluation
  1. 01

    Reason

    Handle difficult multi-step problems with high-default dynamic thinking.

  2. 02

    Ground

    Use search, URLs, code, or other tools when the task needs evidence beyond model memory.

  3. 03

    Execute

    Coordinate tools and intermediate results without losing the original objective.

  4. 04

    Verify

    Check factual consistency and completion quality before returning the final result.

03 / Software engineering

Software Engineering and Vibe-Coding Are Core Use Cases

Google explicitly positions Gemini 3.1 Pro Preview for software-engineering behavior, usability, agentic coding, and vibe-coding workflows.

Judge Whether the Model Can Complete the Engineering Task

A strong coding model must do more than generate syntactically valid code.

It should understand a repository, preserve architecture, inspect dependencies, select relevant files, edit precisely, call tools when necessary, interpret tests, recover from failures, and stop when the implementation is actually correct.

That makes Gemini 3.1 Pro Preview relevant to:

  • repository-level implementation;
  • multi-file refactoring;
  • debugging;
  • architecture changes;
  • code search;
  • test-driven repair loops;
  • terminal or bash agents;
  • rapid application prototyping;
  • UI and application generation.

“Vibe-coding” is useful as a product category, but production evaluation should still be rigorous.

Measure build success, test pass rate, changed-file precision, regression rate, repair-loop count, tool calls, and human corrections.

A model that produces impressive first-pass code but requires repeated manual cleanup may be less useful than a model with slightly slower but more reliable completion.

Engineering workflows
  1. 01

    Repository understanding

    Reason over code, dependencies, documentation, tests, and project structure inside a large working context.

  2. 02

    Implementation

    Make coordinated multi-file changes rather than producing isolated snippets.

  3. 03

    Tool-assisted validation

    Search code, execute commands, run code, and inspect results before declaring success.

  4. 04

    Iterative repair

    Use failed tests and tool output to revise the implementation while preserving task context.

04 / Custom tools

A Separate Endpoint for Bash and Custom-Tool Agents

Gemini 3.1 Pro Preview has an unusual integration option: a dedicated gemini-3.1-pro-preview-customtools endpoint for agents that mix bash-style execution with application-defined tools.

Use It When the Default Model Over-Prioritizes Bash

Google explains that the custom-tools endpoint is better at prioritizing developer-defined tools such as view_file or search_code.

This matters in coding agents.

A model with access to bash may sometimes solve everything through shell commands even when the application provides safer, structured tools with clearer permissions and better observability.

The custom-tools endpoint gives developers another routing option when tool priority becomes part of agent quality.

It uses the same Gemini 3.1 Pro Preview pricing.

However, Google also warns that this endpoint is optimized specifically for workflows that benefit from custom tools plus bash. Quality can fluctuate in tasks that do not need that setup.

Treat the two endpoints as separate evaluation targets rather than assuming the custom-tools variant is universally better.

Dedicated custom-tools endpoint
  1. 01

    Default endpoint

    Use gemini-3.1-pro-preview as the general Pro preview for reasoning, coding, multimodal work, and tools.

  2. 02

    Custom-tools endpoint

    Use gemini-3.1-pro-preview-customtools when the agent mixes bash with tools such as view_file or search_code.

  3. 03

    Measure tool preference

    Track whether the model chooses the intended structured tools instead of unnecessary shell execution.

  4. 04

    Keep routing configurable

    The specialized endpoint can fluctuate on workloads that do not benefit from custom tools, so route by evaluated task class.

05 / Pricing

Gemini 3.1 Pro Preview Pricing

Gemini 3.1 Pro Preview uses two standard token-price tiers based on prompt size, with higher rates once the request exceeds 200,000 prompt tokens.

The 200K Threshold Changes the Economics of Long Context

For prompts up to and including 200K tokens, Google lists:

  • $2.00 per 1M input tokens.
  • $12.00 per 1M output tokens, including thinking tokens.
  • $0.20 per 1M cached-context tokens.

For prompts above 200K tokens:

  • $4.00 per 1M input tokens.
  • $18.00 per 1M output tokens, including thinking tokens.
  • $0.40 per 1M cached-context tokens.

Context-cache storage costs $4.50 per 1M cached tokens per hour.

Batch and Flex inference cost half the Standard token rates:

  • $1 input / $6 output at or below 200K;
  • $2 input / $9 output above 200K.

Priority inference costs more in exchange for a higher-priority serving tier.

This pricing structure is critical for long-context applications.

A 1M-token context window does not mean a 700K-token request has the same unit economics as a 100K-token request. Once the prompt crosses 200K, the entire request enters the higher price tier.

Long-context price tiers

1M tokens · USD

Input
$2.00
Cached input
$0.20
Output + thinking
$12.00

Example: 20K input + 4K output

Input cost
$0.0400
Output cost
$0.0480
Estimated total
$0.0880

06 / Multimodal context

1M Tokens Across Text, Images, Video, Audio, and PDFs

Gemini 3.1 Pro Preview supports up to 1,048,576 input tokens across five input types and up to 65,536 tokens of text output.

Pro-Tier Reasoning Over a Mixed Working Set

The value of a 1M context window is not simply the ability to paste a huge document.

A single difficult workflow can combine:

  • source code;
  • architecture documents;
  • screenshots;
  • long PDFs;
  • recorded audio;
  • video;
  • retrieved URLs;
  • tool results;
  • conversation history;
  • system instructions.

Gemini can reason across that mixed context without requiring each modality to be handled by a separate model.

This is useful for repository audits, multimodal research, technical investigations, complex document analysis, and agent workflows with substantial persistent state.

But the long-context price threshold matters.

Once a prompt exceeds 200K tokens, standard input, output, and cache-read rates increase. Retrieval, selective history, caching, summarization, and context pruning can therefore improve both relevance and economics even when the entire workload technically fits.

Context capacity

Input context

1,048,576

Max output

65,536

Multimodal inputText output

Requests above 200K prompt tokens move into a higher standard pricing tier even though the technical input limit remains 1,048,576 tokens.

07 / Thinking

High-Default Dynamic Thinking

Gemini 3.1 Pro Preview uses dynamic thinking by default at the high level, reflecting its role as the reasoning-focused Pro tier rather than a latency-first model.

Low, Medium, and High Are Supported

Google documents three supported thinking levels:

  • low;
  • medium;
  • high.

high is the default and uses dynamic thinking, allowing the model to adjust reasoning depth within the high allowance based on task complexity.

medium provides a balanced reasoning setting.

low reduces latency and cost for simpler instruction-following or high-throughput requests.

minimal is not supported on Gemini 3.1 Pro Preview.

Google also retains thinking_budget for backward compatibility, but recommends thinking_level for more predictable performance. The two parameters cannot be used in the same request.

Thinking tokens are billed as output tokens.

For a Pro model with $12 or $18 per-million output pricing, reasoning configuration can materially change request economics.

Thinking levels
lowmediumhigh · default

Simpler request

Use low or medium when the task does not need maximum reasoning depth and latency or output cost matters.

Difficult Pro workload

Keep high for complex multimodal reasoning, difficult software engineering, and tool-heavy workflows where deeper reasoning improves accepted-result quality.

08 / Tool ecosystem

A Broad Tool Surface for Grounded Pro Workflows

Gemini 3.1 Pro Preview combines high-level reasoning with search, maps, code execution, URLs, custom functions, structured outputs, caching, and file retrieval in supported environments.

Tool Composition Is a Core Gemini 3 Feature

Google lists support for:

  • Google Search grounding;
  • Google Maps grounding;
  • code execution;
  • function calling;
  • structured outputs;
  • URL context;
  • context caching;
  • File Search in AI Studio;
  • Batch API;
  • Flex inference;
  • Priority inference.

Gemini 3 also supports combining built-in and custom tools in a single workflow.

That is important for agents that need both external information and application actions.

For example, the model can search the web, inspect a URL, execute code, and then return schema-constrained structured output.

Structured outputs can also be combined with tools so grounded or computed results remain machine-readable.

The base Gemini 3.1 Pro Preview model does not support audio generation, image generation, or the Live API.

Supported capabilities
  • Google Search

    Ground answers in current web information.

    Supported
  • Google Maps

    Use Maps grounding in location-aware workflows.

    Supported
  • Code execution

    Run code for calculations, data processing, and verification.

    Supported
  • Function calling

    Invoke application-defined tools and actions.

    Supported
  • Structured outputs

    Return schema-constrained machine-readable responses, including with supported tools.

    Supported
  • URL context

    Read and reason over content from supplied URLs.

    Supported
  • Context caching

    Reuse large repeated prompt prefixes with a 4,096-token minimum for implicit caching.

    Supported
  • Live API

    The model page does not list Live API support for Gemini 3.1 Pro Preview.

    Not listed

09 / Preview lifecycle

Preview Status Is a Production Architecture Constraint

Gemini 3.1 Pro Preview remains a public-preview endpoint, so applications should separate model evaluation from long-term dependency planning.

Do Not Treat Preview Like a Permanent Stable Contract

Google released the model in February 2026 and currently lists no announced shutdown date.

That does not make it equivalent to a stable model.

Preview models can change faster, be replaced by newer versions, or receive different lifecycle treatment than stable endpoints.

The previous Gemini 3 Pro Preview demonstrates why this matters: Google shut that model down on March 9, 2026 and redirected the alias to Gemini 3.1 Pro Preview.

For production architecture:

  • keep the model ID configurable;
  • maintain regression prompts;
  • isolate provider-specific thinking and tool settings;
  • test a fallback or successor;
  • monitor deprecation announcements;
  • avoid hard-coding preview assumptions into business logic.

Preview can be the right choice when its capability materially improves the product. It should not be an excuse to make migration expensive.

Preview deployment checklist
  1. 01

    Keep routing configurable

    Treat gemini-3.1-pro-preview as configuration rather than embedding the preview ID throughout application logic.

  2. 02

    Maintain regression evals

    Keep representative prompts, tool trajectories, multimodal inputs, and acceptance criteria ready for successor testing.

  3. 03

    Watch lifecycle changes

    Monitor official deprecation and release notes because preview endpoints can be replaced before stable models.

  4. 04

    Design a fallback

    Know which stable or newer model can absorb critical traffic if preview availability or behavior changes.

10 / 3 Pro → 3.1 Pro

Migrating From Gemini 3 Pro Preview to 3.1 Pro Preview

Gemini 3.1 Pro Preview replaced the original Gemini 3 Pro Preview and preserves the 1M/64K Pro envelope while expanding reasoning controls, tooling, serving options, and reliability.

The Old Preview Is Already Shut Down

Google shut down gemini-3-pro-preview on March 9, 2026.

The old gemini-3-pro-preview alias now points to gemini-3.1-pro-preview.

Both models use a 1,048,576-token input limit and 65,536-token output limit, but 3.1 Pro adds or improves several practical areas.

Google positions 3.1 Pro around better thinking, token efficiency, factual consistency, software-engineering usability, precise tool usage, and reliable multi-step execution.

The old Gemini 3 Pro Preview supported only low and high thinking levels. Gemini 3.1 Pro Preview also supports medium.

3.1 Pro additionally supports Google Maps grounding, Flex inference, and Priority inference in its current model specification.

It also introduces the dedicated custom-tools endpoint for bash-heavy agents.

Migration improvements
  1. 01

    Move to 3.1 Pro

    The original gemini-3-pro-preview endpoint was shut down on March 9, 2026 and its alias now resolves to 3.1 Pro Preview.

  2. 02

    Re-evaluate thinking

    3.1 Pro adds medium between low and high, while high remains the default dynamic setting.

  3. 03

    Expand tool and serving options

    Review Maps grounding, Flex, Priority, combined tools, and the specialized custom-tools endpoint.

  4. 04

    Replay reliability failures

    Use tasks where Gemini 3 Pro produced factual inconsistency, poor tool choice, or multi-step failures to validate the 3.1 upgrade.

11 / Evaluation

Gemini 3.1 Pro Preview Strengths and Limitations

Gemini 3.1 Pro Preview is best evaluated as a high-capability reasoning and agent model: stronger than Flash for difficult workflows, but more expensive, high-thinking by default, and still subject to preview lifecycle risk.

Strengths

  • Pro-tier reasoning

    High-default dynamic thinking is designed for complex reasoning across text, images, video, audio, PDFs, and tool results.

  • Software-engineering focus

    Google explicitly optimizes the model for software-engineering behavior, usability, vibe-coding, and reliable multi-step execution.

  • Dedicated custom-tools path

    The customtools endpoint provides a specialized option for agents combining bash with developer-defined tools.

  • Broad grounded tool ecosystem

    Search, Maps, code execution, URLs, structured outputs, functions, caching, Batch, Flex, and Priority support sophisticated production workflows.

What to consider

  • Preview lifecycle

    The endpoint is public preview with no announced shutdown date, so production systems should maintain migration and fallback paths.

  • Higher Pro pricing

    Standard pricing starts at $2/$12 and rises to $4/$18 when the prompt exceeds 200K tokens.

  • High thinking by default

    Difficult reasoning can increase latency and output-token cost; simpler workloads may be better served with medium, low, or a Flash model.

  • Not every Gemini surface is supported

    Audio generation, image generation, and Live API support are not listed for this model, and File Search support is environment-specific.

Evaluate the Pro tier with real evidence

Test Gemini 3.1 Pro Preview on your hardest workflows

Replay complex reasoning, software engineering, multimodal analysis, search-grounded research, and custom-tool agents to compare accuracy, thinking cost, tool-call reliability, latency, and total task completion.

Start Free

This is a preview model. Benchmark it aggressively, but keep model routing and migration paths configurable because preview endpoints can change before a stable release.

Common Questions

What is Gemini 3.1 Pro Preview?

Gemini 3.1 Pro Preview is Google’s public-preview Pro model for complex multimodal reasoning, software engineering, vibe-coding, and agentic workflows that require precise tool use and reliable multi-step execution.

What is the Gemini 3.1 Pro Preview model ID?

The primary model ID is gemini-3.1-pro-preview. Google also offers gemini-3.1-pro-preview-customtools for agents that mix bash with developer-defined tools.

When was Gemini 3.1 Pro Preview released?

Google released Gemini 3.1 Pro Preview on February 19, 2026.

Is Gemini 3.1 Pro Preview stable?

No. Google lists the model as Public preview. No shutdown date is currently announced, but production applications should keep migration and fallback paths configurable.

What is the Gemini 3.1 Pro Preview context window?

Gemini 3.1 Pro Preview supports up to 1,048,576 input tokens and up to 65,536 output tokens.

What input types does Gemini 3.1 Pro Preview support?

The model accepts text, images, video, audio, and PDF input and produces text output.

What is the Gemini 3.1 Pro Preview knowledge cutoff?

Google documents January 2025 as the Gemini 3 knowledge cutoff. Google Search grounding can be used when current external information is required.

How much does Gemini 3.1 Pro Preview cost?

For prompts up to 200K tokens, standard pricing is $2 per 1M input tokens and $12 per 1M output tokens, including thinking. For prompts above 200K tokens, pricing increases to $4 input and $18 output per million tokens.

How much does Gemini 3.1 Pro Preview context caching cost?

Standard cache-read pricing is $0.20 per 1M tokens for prompts up to 200K and $0.40 above 200K. Google also lists $4.50 per 1M cached tokens per hour for storage.

Does Gemini 3.1 Pro Preview have a free Gemini API tier?

No. Google documents no free Gemini API tier for gemini-3.1-pro-preview, although the model can be tried in Google AI Studio.

Does Gemini 3.1 Pro Preview support thinking?

Yes. It supports low, medium, and high thinking levels, with high as the default dynamic setting.

Does Gemini 3.1 Pro Preview support minimal thinking?

No. The supported levels are low, medium, and high. Use low when you need to reduce reasoning latency and cost.

Can I still use thinking_budget with Gemini 3.1 Pro Preview?

Google retains thinking_budget for backward compatibility but recommends thinking_level. Do not send both in the same request because the API returns a 400 error.

What is gemini-3.1-pro-preview-customtools?

It is a separate Gemini 3.1 Pro Preview endpoint optimized for agentic workflows that combine bash with custom tools such as view_file or search_code. Google notes that it can prioritize custom tools more effectively.

Does the custom-tools endpoint cost more?

No. Google lists the Gemini 3.1 Pro Preview Custom Tools endpoint at the same token pricing as the standard Gemini 3.1 Pro Preview endpoint.

What tools does Gemini 3.1 Pro Preview support?

Google lists Search grounding, Maps grounding, code execution, function calling, structured outputs, URL context, caching, and File Search in AI Studio. Batch, Flex, and Priority inference are also supported.

Does Gemini 3.1 Pro Preview support the Live API?

No. Google’s model specification currently lists Live API as unsupported for Gemini 3.1 Pro Preview.

How is Gemini 3.1 Pro Preview different from Gemini 3 Pro Preview?

Gemini 3.1 Pro Preview replaces the shut-down Gemini 3 Pro Preview and adds stronger thinking and reliability positioning, medium thinking, Maps grounding, Flex and Priority inference, and a dedicated custom-tools endpoint while preserving the 1M input and 64K output envelope.

How is Gemini 3.1 Pro Preview different from Gemini 3.8 Flash?

Gemini 3.1 Pro Preview is the higher-priced Pro preview for difficult reasoning and tool-heavy workflows, with high thinking by default and long-context pricing tiers. Gemini 3.8 Flash is the lower-cost stable Flash workhorse optimized for long-horizon engineering and autonomous agents at higher throughput.

Should I use Gemini 3.1 Pro Preview for a new application?

Use it when Pro-tier reasoning, multimodal analysis, coding quality, or precise tool orchestration materially improves your workload. Because the endpoint remains in preview, keep model routing configurable and benchmark stable Flash alternatives for fallback and cost-sensitive traffic.

Model information

Last updated

Specifications, pricing, preview status, thinking levels, multimodal input types, software-engineering positioning, custom-tools endpoint behavior, caching, inference options, and migration guidance on this page are based on official Google Gemini API and Google Cloud documentation.

Gemini 3.1 Pro Preview — Pricing, 1M Context, Reasoning & Custom Tools | EidoStack