Google Gemini preview model
Gemini 3.1 Pro Preview
Google’s preview Pro-tier Gemini model for complex multimodal reasoning, software engineering, vibe-coding, and agentic workflows that need precise tool use and reliable multi-step execution.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $2.00
- ≤200K prompt · per 1M
- Cached input
- $0.20
- ≤200K prompt · per 1M
- Output
- $12.00
- ≤200K prompt · per 1M
01 / Overview
What Gemini 3.1 Pro Preview Is
Gemini 3.1 Pro Preview is Google’s preview Pro-tier model for complex tasks that need broad world knowledge, advanced multimodal reasoning, software-engineering capability, and reliable agentic execution.
A Pro Model Focused on Reliability, Not Throughput
Google describes Gemini 3.1 Pro Preview as a refinement of the Gemini 3 Pro series with better thinking, improved token efficiency, and a more grounded, factually consistent experience.
Its product role is therefore different from Gemini Flash.
Flash models optimize the capability-throughput-cost balance. Gemini 3.1 Pro Preview is the model to evaluate when the task is difficult enough that reasoning quality, factual consistency, software-engineering behavior, and precise tool usage matter more than lowest token price.
Google released gemini-3.1-pro-preview on February 19, 2026.
The model accepts text, images, video, audio, and PDFs and returns text. It supports up to 1,048,576 input tokens and up to 65,536 output tokens.
It also supports dynamic thinking, Google Search grounding, Google Maps grounding, code execution, function calling, structured outputs, URL context, context caching, and several inference tiers.
- Model ID:
gemini-3.1-pro-preview. - Launch stage: Public preview.
- Released February 19, 2026.
- 1,048,576 input tokens.
- 65,536 maximum output tokens.
- Text, image, video, audio, and PDF input.
- Text output.
- Default thinking: high.
- Separate custom-tools endpoint available.
- Provider
- Family
- Gemini 3.1 Pro
- Model ID
- gemini-3.1-pro-preview
- Launch stage
- Public preview
- Knowledge cutoff
- Jan 2025
- Input
- Text, Image, Video, Audio, PDF
- Output
- Text
02 / Reasoning & reliability
Better Thinking and More Grounded Multi-Step Execution
Google positions Gemini 3.1 Pro Preview around better thinking, improved token efficiency, stronger grounding, factual consistency, and more reliable execution across complex real-world tasks.
Evaluate the Entire Reasoning Trajectory
A difficult production task is rarely one prompt followed by one answer.
The model may need to understand a large context, form a plan, search for external evidence, call tools, evaluate results, revise its approach, and produce a final answer that remains consistent with everything it observed.
Failures can happen at any point:
- the initial plan can be wrong;
- a tool can be selected incorrectly;
- valid evidence can be ignored;
- a later answer can contradict an earlier observation;
- the model can spend too many tokens before reaching a useful result.
Gemini 3.1 Pro Preview is specifically intended to improve this class of behavior.
For evaluation, measure:
- final-answer correctness;
- factual consistency across long responses;
- successful tool selection;
- valid tool arguments;
- recovery after failed calls;
- number of unnecessary steps;
- thinking-token usage;
- time to accepted result.
The relevant unit is not “one completion.” It is the complete reasoning-and-action trajectory.
- 01
Reason
Handle difficult multi-step problems with high-default dynamic thinking.
- 02
Ground
Use search, URLs, code, or other tools when the task needs evidence beyond model memory.
- 03
Execute
Coordinate tools and intermediate results without losing the original objective.
- 04
Verify
Check factual consistency and completion quality before returning the final result.
03 / Software engineering
Software Engineering and Vibe-Coding Are Core Use Cases
Google explicitly positions Gemini 3.1 Pro Preview for software-engineering behavior, usability, agentic coding, and vibe-coding workflows.
Judge Whether the Model Can Complete the Engineering Task
A strong coding model must do more than generate syntactically valid code.
It should understand a repository, preserve architecture, inspect dependencies, select relevant files, edit precisely, call tools when necessary, interpret tests, recover from failures, and stop when the implementation is actually correct.
That makes Gemini 3.1 Pro Preview relevant to:
- repository-level implementation;
- multi-file refactoring;
- debugging;
- architecture changes;
- code search;
- test-driven repair loops;
- terminal or bash agents;
- rapid application prototyping;
- UI and application generation.
“Vibe-coding” is useful as a product category, but production evaluation should still be rigorous.
Measure build success, test pass rate, changed-file precision, regression rate, repair-loop count, tool calls, and human corrections.
A model that produces impressive first-pass code but requires repeated manual cleanup may be less useful than a model with slightly slower but more reliable completion.
- 01
Repository understanding
Reason over code, dependencies, documentation, tests, and project structure inside a large working context.
- 02
Implementation
Make coordinated multi-file changes rather than producing isolated snippets.
- 03
Tool-assisted validation
Search code, execute commands, run code, and inspect results before declaring success.
- 04
Iterative repair
Use failed tests and tool output to revise the implementation while preserving task context.
04 / Custom tools
A Separate Endpoint for Bash and Custom-Tool Agents
Gemini 3.1 Pro Preview has an unusual integration option: a dedicated gemini-3.1-pro-preview-customtools endpoint for agents that mix bash-style execution with application-defined tools.
Use It When the Default Model Over-Prioritizes Bash
Google explains that the custom-tools endpoint is better at prioritizing developer-defined tools such as view_file or search_code.
This matters in coding agents.
A model with access to bash may sometimes solve everything through shell commands even when the application provides safer, structured tools with clearer permissions and better observability.
The custom-tools endpoint gives developers another routing option when tool priority becomes part of agent quality.
It uses the same Gemini 3.1 Pro Preview pricing.
However, Google also warns that this endpoint is optimized specifically for workflows that benefit from custom tools plus bash. Quality can fluctuate in tasks that do not need that setup.
Treat the two endpoints as separate evaluation targets rather than assuming the custom-tools variant is universally better.
- 01
Default endpoint
Use gemini-3.1-pro-preview as the general Pro preview for reasoning, coding, multimodal work, and tools.
- 02
Custom-tools endpoint
Use gemini-3.1-pro-preview-customtools when the agent mixes bash with tools such as view_file or search_code.
- 03
Measure tool preference
Track whether the model chooses the intended structured tools instead of unnecessary shell execution.
- 04
Keep routing configurable
The specialized endpoint can fluctuate on workloads that do not benefit from custom tools, so route by evaluated task class.
05 / Pricing
Gemini 3.1 Pro Preview Pricing
Gemini 3.1 Pro Preview uses two standard token-price tiers based on prompt size, with higher rates once the request exceeds 200,000 prompt tokens.
The 200K Threshold Changes the Economics of Long Context
For prompts up to and including 200K tokens, Google lists:
- $2.00 per 1M input tokens.
- $12.00 per 1M output tokens, including thinking tokens.
- $0.20 per 1M cached-context tokens.
For prompts above 200K tokens:
- $4.00 per 1M input tokens.
- $18.00 per 1M output tokens, including thinking tokens.
- $0.40 per 1M cached-context tokens.
Context-cache storage costs $4.50 per 1M cached tokens per hour.
Batch and Flex inference cost half the Standard token rates:
- $1 input / $6 output at or below 200K;
- $2 input / $9 output above 200K.
Priority inference costs more in exchange for a higher-priority serving tier.
This pricing structure is critical for long-context applications.
A 1M-token context window does not mean a 700K-token request has the same unit economics as a 100K-token request. Once the prompt crosses 200K, the entire request enters the higher price tier.
1M tokens · USD
- Input
- $2.00
- Cached input
- $0.20
- Output + thinking
- $12.00
Example: 20K input + 4K output
- Input cost
- $0.0400
- Output cost
- $0.0480
- Estimated total
- $0.0880
06 / Multimodal context
1M Tokens Across Text, Images, Video, Audio, and PDFs
Gemini 3.1 Pro Preview supports up to 1,048,576 input tokens across five input types and up to 65,536 tokens of text output.
Pro-Tier Reasoning Over a Mixed Working Set
The value of a 1M context window is not simply the ability to paste a huge document.
A single difficult workflow can combine:
- source code;
- architecture documents;
- screenshots;
- long PDFs;
- recorded audio;
- video;
- retrieved URLs;
- tool results;
- conversation history;
- system instructions.
Gemini can reason across that mixed context without requiring each modality to be handled by a separate model.
This is useful for repository audits, multimodal research, technical investigations, complex document analysis, and agent workflows with substantial persistent state.
But the long-context price threshold matters.
Once a prompt exceeds 200K tokens, standard input, output, and cache-read rates increase. Retrieval, selective history, caching, summarization, and context pruning can therefore improve both relevance and economics even when the entire workload technically fits.
Input context
1,048,576
Max output
65,536
Requests above 200K prompt tokens move into a higher standard pricing tier even though the technical input limit remains 1,048,576 tokens.
07 / Thinking
High-Default Dynamic Thinking
Gemini 3.1 Pro Preview uses dynamic thinking by default at the high level, reflecting its role as the reasoning-focused Pro tier rather than a latency-first model.
Low, Medium, and High Are Supported
Google documents three supported thinking levels:
low;medium;high.
high is the default and uses dynamic thinking, allowing the model to adjust reasoning depth within the high allowance based on task complexity.
medium provides a balanced reasoning setting.
low reduces latency and cost for simpler instruction-following or high-throughput requests.
minimal is not supported on Gemini 3.1 Pro Preview.
Google also retains thinking_budget for backward compatibility, but recommends thinking_level for more predictable performance. The two parameters cannot be used in the same request.
Thinking tokens are billed as output tokens.
For a Pro model with $12 or $18 per-million output pricing, reasoning configuration can materially change request economics.
Simpler request
Use low or medium when the task does not need maximum reasoning depth and latency or output cost matters.
Difficult Pro workload
Keep high for complex multimodal reasoning, difficult software engineering, and tool-heavy workflows where deeper reasoning improves accepted-result quality.
08 / Tool ecosystem
A Broad Tool Surface for Grounded Pro Workflows
Gemini 3.1 Pro Preview combines high-level reasoning with search, maps, code execution, URLs, custom functions, structured outputs, caching, and file retrieval in supported environments.
Tool Composition Is a Core Gemini 3 Feature
Google lists support for:
- Google Search grounding;
- Google Maps grounding;
- code execution;
- function calling;
- structured outputs;
- URL context;
- context caching;
- File Search in AI Studio;
- Batch API;
- Flex inference;
- Priority inference.
Gemini 3 also supports combining built-in and custom tools in a single workflow.
That is important for agents that need both external information and application actions.
For example, the model can search the web, inspect a URL, execute code, and then return schema-constrained structured output.
Structured outputs can also be combined with tools so grounded or computed results remain machine-readable.
The base Gemini 3.1 Pro Preview model does not support audio generation, image generation, or the Live API.
- Supported
Google Search
Ground answers in current web information.
- Supported
Google Maps
Use Maps grounding in location-aware workflows.
- Supported
Code execution
Run code for calculations, data processing, and verification.
- Supported
Function calling
Invoke application-defined tools and actions.
- Supported
Structured outputs
Return schema-constrained machine-readable responses, including with supported tools.
- Supported
URL context
Read and reason over content from supplied URLs.
- Supported
Context caching
Reuse large repeated prompt prefixes with a 4,096-token minimum for implicit caching.
- Not listed
Live API
The model page does not list Live API support for Gemini 3.1 Pro Preview.
09 / Preview lifecycle
Preview Status Is a Production Architecture Constraint
Gemini 3.1 Pro Preview remains a public-preview endpoint, so applications should separate model evaluation from long-term dependency planning.
Do Not Treat Preview Like a Permanent Stable Contract
Google released the model in February 2026 and currently lists no announced shutdown date.
That does not make it equivalent to a stable model.
Preview models can change faster, be replaced by newer versions, or receive different lifecycle treatment than stable endpoints.
The previous Gemini 3 Pro Preview demonstrates why this matters: Google shut that model down on March 9, 2026 and redirected the alias to Gemini 3.1 Pro Preview.
For production architecture:
- keep the model ID configurable;
- maintain regression prompts;
- isolate provider-specific thinking and tool settings;
- test a fallback or successor;
- monitor deprecation announcements;
- avoid hard-coding preview assumptions into business logic.
Preview can be the right choice when its capability materially improves the product. It should not be an excuse to make migration expensive.
- 01
Keep routing configurable
Treat gemini-3.1-pro-preview as configuration rather than embedding the preview ID throughout application logic.
- 02
Maintain regression evals
Keep representative prompts, tool trajectories, multimodal inputs, and acceptance criteria ready for successor testing.
- 03
Watch lifecycle changes
Monitor official deprecation and release notes because preview endpoints can be replaced before stable models.
- 04
Design a fallback
Know which stable or newer model can absorb critical traffic if preview availability or behavior changes.
10 / 3 Pro → 3.1 Pro
Migrating From Gemini 3 Pro Preview to 3.1 Pro Preview
Gemini 3.1 Pro Preview replaced the original Gemini 3 Pro Preview and preserves the 1M/64K Pro envelope while expanding reasoning controls, tooling, serving options, and reliability.
The Old Preview Is Already Shut Down
Google shut down gemini-3-pro-preview on March 9, 2026.
The old gemini-3-pro-preview alias now points to gemini-3.1-pro-preview.
Both models use a 1,048,576-token input limit and 65,536-token output limit, but 3.1 Pro adds or improves several practical areas.
Google positions 3.1 Pro around better thinking, token efficiency, factual consistency, software-engineering usability, precise tool usage, and reliable multi-step execution.
The old Gemini 3 Pro Preview supported only low and high thinking levels. Gemini 3.1 Pro Preview also supports medium.
3.1 Pro additionally supports Google Maps grounding, Flex inference, and Priority inference in its current model specification.
It also introduces the dedicated custom-tools endpoint for bash-heavy agents.
- 01
Move to 3.1 Pro
The original gemini-3-pro-preview endpoint was shut down on March 9, 2026 and its alias now resolves to 3.1 Pro Preview.
- 02
Re-evaluate thinking
3.1 Pro adds medium between low and high, while high remains the default dynamic setting.
- 03
Expand tool and serving options
Review Maps grounding, Flex, Priority, combined tools, and the specialized custom-tools endpoint.
- 04
Replay reliability failures
Use tasks where Gemini 3 Pro produced factual inconsistency, poor tool choice, or multi-step failures to validate the 3.1 upgrade.
11 / Evaluation
Gemini 3.1 Pro Preview Strengths and Limitations
Gemini 3.1 Pro Preview is best evaluated as a high-capability reasoning and agent model: stronger than Flash for difficult workflows, but more expensive, high-thinking by default, and still subject to preview lifecycle risk.
Strengths
Pro-tier reasoning
High-default dynamic thinking is designed for complex reasoning across text, images, video, audio, PDFs, and tool results.
Software-engineering focus
Google explicitly optimizes the model for software-engineering behavior, usability, vibe-coding, and reliable multi-step execution.
Dedicated custom-tools path
The customtools endpoint provides a specialized option for agents combining bash with developer-defined tools.
Broad grounded tool ecosystem
Search, Maps, code execution, URLs, structured outputs, functions, caching, Batch, Flex, and Priority support sophisticated production workflows.
What to consider
Preview lifecycle
The endpoint is public preview with no announced shutdown date, so production systems should maintain migration and fallback paths.
Higher Pro pricing
Standard pricing starts at $2/$12 and rises to $4/$18 when the prompt exceeds 200K tokens.
High thinking by default
Difficult reasoning can increase latency and output-token cost; simpler workloads may be better served with medium, low, or a Flash model.
Not every Gemini surface is supported
Audio generation, image generation, and Live API support are not listed for this model, and File Search support is environment-specific.
Evaluate the Pro tier with real evidence
Test Gemini 3.1 Pro Preview on your hardest workflows
Replay complex reasoning, software engineering, multimodal analysis, search-grounded research, and custom-tool agents to compare accuracy, thinking cost, tool-call reliability, latency, and total task completion.
Start FreeThis is a preview model. Benchmark it aggressively, but keep model routing and migration paths configurable because preview endpoints can change before a stable release.
Common Questions
What is Gemini 3.1 Pro Preview?
Gemini 3.1 Pro Preview is Google’s public-preview Pro model for complex multimodal reasoning, software engineering, vibe-coding, and agentic workflows that require precise tool use and reliable multi-step execution.
What is the Gemini 3.1 Pro Preview model ID?
The primary model ID is gemini-3.1-pro-preview. Google also offers gemini-3.1-pro-preview-customtools for agents that mix bash with developer-defined tools.
When was Gemini 3.1 Pro Preview released?
Google released Gemini 3.1 Pro Preview on February 19, 2026.
Is Gemini 3.1 Pro Preview stable?
No. Google lists the model as Public preview. No shutdown date is currently announced, but production applications should keep migration and fallback paths configurable.
What is the Gemini 3.1 Pro Preview context window?
Gemini 3.1 Pro Preview supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 3.1 Pro Preview support?
The model accepts text, images, video, audio, and PDF input and produces text output.
What is the Gemini 3.1 Pro Preview knowledge cutoff?
Google documents January 2025 as the Gemini 3 knowledge cutoff. Google Search grounding can be used when current external information is required.
How much does Gemini 3.1 Pro Preview cost?
For prompts up to 200K tokens, standard pricing is $2 per 1M input tokens and $12 per 1M output tokens, including thinking. For prompts above 200K tokens, pricing increases to $4 input and $18 output per million tokens.
How much does Gemini 3.1 Pro Preview context caching cost?
Standard cache-read pricing is $0.20 per 1M tokens for prompts up to 200K and $0.40 above 200K. Google also lists $4.50 per 1M cached tokens per hour for storage.
Does Gemini 3.1 Pro Preview have a free Gemini API tier?
No. Google documents no free Gemini API tier for gemini-3.1-pro-preview, although the model can be tried in Google AI Studio.
Does Gemini 3.1 Pro Preview support thinking?
Yes. It supports low, medium, and high thinking levels, with high as the default dynamic setting.
Does Gemini 3.1 Pro Preview support minimal thinking?
No. The supported levels are low, medium, and high. Use low when you need to reduce reasoning latency and cost.
Can I still use thinking_budget with Gemini 3.1 Pro Preview?
Google retains thinking_budget for backward compatibility but recommends thinking_level. Do not send both in the same request because the API returns a 400 error.
What is gemini-3.1-pro-preview-customtools?
It is a separate Gemini 3.1 Pro Preview endpoint optimized for agentic workflows that combine bash with custom tools such as view_file or search_code. Google notes that it can prioritize custom tools more effectively.
Does the custom-tools endpoint cost more?
No. Google lists the Gemini 3.1 Pro Preview Custom Tools endpoint at the same token pricing as the standard Gemini 3.1 Pro Preview endpoint.
What tools does Gemini 3.1 Pro Preview support?
Google lists Search grounding, Maps grounding, code execution, function calling, structured outputs, URL context, caching, and File Search in AI Studio. Batch, Flex, and Priority inference are also supported.
Does Gemini 3.1 Pro Preview support the Live API?
No. Google’s model specification currently lists Live API as unsupported for Gemini 3.1 Pro Preview.
How is Gemini 3.1 Pro Preview different from Gemini 3 Pro Preview?
Gemini 3.1 Pro Preview replaces the shut-down Gemini 3 Pro Preview and adds stronger thinking and reliability positioning, medium thinking, Maps grounding, Flex and Priority inference, and a dedicated custom-tools endpoint while preserving the 1M input and 64K output envelope.
How is Gemini 3.1 Pro Preview different from Gemini 3.8 Flash?
Gemini 3.1 Pro Preview is the higher-priced Pro preview for difficult reasoning and tool-heavy workflows, with high thinking by default and long-context pricing tiers. Gemini 3.8 Flash is the lower-cost stable Flash workhorse optimized for long-horizon engineering and autonomous agents at higher throughput.
Should I use Gemini 3.1 Pro Preview for a new application?
Use it when Pro-tier reasoning, multimodal analysis, coding quality, or precise tool orchestration materially improves your workload. Because the endpoint remains in preview, keep model routing configurable and benchmark stable Flash alternatives for fallback and cost-sensitive traffic.
Model information
Last updated
Specifications, pricing, preview status, thinking levels, multimodal input types, software-engineering positioning, custom-tools endpoint behavior, caching, inference options, and migration guidance on this page are based on official Google Gemini API and Google Cloud documentation.