Google Gemini premium reasoning model
Gemini 2.5 Pro
Google’s flagship Gemini 2.5 thinking model for complex coding, mathematics, STEM reasoning, large codebases, multimodal analysis, and long-context workloads where quality matters more than Flash-tier latency.
- Input context
- 1.05M
- tokens
- Max output
- 65.5K
- tokens
- Input
- $1.25
- ≤200K · per 1M
- Cached input
- $0.125
- ≤200K · per 1M
- Output
- $10.00
- ≤200K · per 1M
01 / Overview
What Gemini 2.5 Pro Is
Gemini 2.5 Pro is Google’s flagship thinking model from the Gemini 2.5 generation, built for complex problems in coding, mathematics, science, large datasets, codebases, and long documents.
The Premium Capability Tier of Gemini 2.5
Google first introduced Gemini 2.5 Pro in March 2025 and released the stable gemini-2.5-pro endpoint as generally available on June 17, 2025.
The model’s role is different from Gemini 2.5 Flash and Flash-Lite.
Flash models optimize the quality-speed-cost tradeoff.
Pro is designed for the workloads where the value of a stronger result can justify more reasoning, more latency, and higher token pricing.
Google describes Gemini 2.5 Pro as a state-of-the-art thinking model for:
- complex coding;
- mathematics;
- STEM reasoning;
- large datasets;
- large codebases;
- long documents;
- multimodal analysis.
It accepts text, images, video, audio, and PDFs and produces text.
The stable endpoint supports:
- 1,048,576 input tokens;
- 65,536 output tokens;
- always-on thinking;
- context caching;
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- function calling;
- structured outputs;
- URL context;
- Batch, Flex, and Priority inference.
- Provider
- Family
- Gemini 2.5 Pro
- Model ID
- gemini-2.5-pro
- Status
- Stable · legacy access
- GA date
- Jun 17, 2025
- Knowledge cutoff
- Jan 2025
- Output
- Text
02 / Deep reasoning
Built for Problems Where Deeper Reasoning Changes the Answer
Gemini 2.5 Pro is not primarily a throughput model. Its value appears when a task requires several connected reasoning steps and a weaker model would create retries, incomplete analysis, or incorrect conclusions.
The Pro Tier Should Earn Its Higher Cost
Google positioned Gemini 2.5 Pro around difficult reasoning in:
- mathematics;
- science;
- knowledge-intensive tasks;
- coding;
- multi-step problem solving;
- large-context analysis.
This makes it a candidate for problems such as:
- debugging a system across several interacting components;
- analyzing a technical specification and implementation together;
- reasoning over mathematical constraints;
- comparing evidence across long documents;
- planning a non-trivial code change;
- interpreting multimodal evidence before choosing a tool action.
The important production metric is not whether Pro produces a stronger answer in isolation.
It is whether that stronger answer reduces enough failures, retries, human interventions, or downstream errors to justify premium pricing.
Measure:
- first-pass acceptance;
- factual and logical correctness;
- reasoning-token use;
- latency;
- retries;
- human review;
- downstream task success;
- cost per accepted result.
- 01
Decompose
Break difficult problems into connected constraints, evidence, and intermediate decisions.
- 02
Reason
Spend additional internal tokens when complex coding, mathematics, STEM, or analysis benefits from deeper deliberation.
- 03
Verify
Use tools, execution results, retrieved evidence, or structured checks to validate intermediate conclusions.
- 04
Measure the premium
Compare stronger completion against added latency and token cost rather than assuming the Pro tier is always necessary.
03 / Coding & web apps
Coding and Web Development Were Core Gemini 2.5 Pro Strengths
Google repeatedly positioned Gemini 2.5 Pro as its strongest Gemini 2.5 model for coding, code transformation, debugging, web development, and complex agentic engineering workflows.
Evaluate Complete Engineering Trajectories
During 2025, Google specifically improved Gemini 2.5 Pro for building interactive web applications.
The model was also integrated into developer products such as Gemini CLI and Gemini Code Assist.
Useful software-engineering workloads include:
- repository analysis;
- code generation;
- code transformation;
- refactoring;
- debugging;
- UI implementation;
- architecture reasoning;
- test generation;
- complex multi-file changes;
- agentic tool use.
A coding benchmark measures only part of this work.
A production coding assistant may need to read several files, understand the architecture, formulate a patch, call tools, inspect errors, revise the implementation, and explain the result.
Measure:
- tests passed;
- accepted implementation rate;
- number of iterations;
- unnecessary edits;
- tool-call count;
- repeated file reads;
- human corrections;
- latency to accepted result.
- 01
Understand
Analyze large repositories, architecture, dependencies, specifications, and surrounding code before editing.
- 02
Implement
Generate or transform code across multiple files while preserving project constraints.
- 03
Debug
Reason from errors, logs, tests, and execution feedback rather than stopping at the first hypothesis.
- 04
Verify
Evaluate whether the complete implementation passes tests and satisfies acceptance criteria.
04 / Pricing
Gemini 2.5 Pro Pricing
Gemini 2.5 Pro uses tiered pricing: requests above 200,000 input tokens cost more per token across both input and output.
A 1M Context Window Has a 200K Pricing Boundary
For Standard paid inference with prompts up to and including 200K tokens:
- input: $1.25 per 1M tokens;
- output including thinking: $10.00 per 1M tokens;
- cached context: $0.125 per 1M tokens.
For prompts above 200K tokens:
- input: $2.50 per 1M tokens;
- output including thinking: $15.00 per 1M tokens;
- cached context: $0.25 per 1M tokens.
Cache storage is $4.50 per 1M tokens per hour.
Batch and Flex inference cut generation rates in half:
≤200K
- $0.625 input;
- $5.00 output.
>200K
- $1.25 input;
- $7.50 output.
Priority inference is more expensive for higher-priority serving.
This threshold is one of the most important economic properties of Gemini 2.5 Pro.
The model can accept more than one million input tokens, but crossing 200K changes the price of the request.
1M tokens · USD
- Input
- $1.25
- Cached input
- $0.125
- Output + thinking
- $10.00
Example: 100K input + 8K output
- Input cost
- $0.1250
- Output cost
- $0.0800
- Estimated total
- $0.2050
05 / Long context
A 1M-Token Window for Codebases, Documents, and Large Datasets
Gemini 2.5 Pro supports up to 1,048,576 input tokens and up to 65,536 output tokens, making long-context analysis one of its defining capabilities.
Pro Is Designed to Reason Across the Working Set
The context can contain:
- source code;
- repository files;
- long PDFs;
- images;
- video;
- audio;
- retrieved documents;
- URLs;
- conversation history;
- system instructions;
- tool results.
Google specifically describes the model as capable of analyzing large datasets, codebases, and documents using long context.
This makes it useful for:
- repository-wide code analysis;
- technical due diligence;
- legal or policy document review;
- long research packets;
- multimodal evidence analysis;
- large-document question answering;
- cross-document synthesis.
However, more context is not automatically better.
Above 200K input tokens, Gemini 2.5 Pro also enters its higher pricing tier.
For production systems, track both:
- whether additional context improves the answer;
- whether that improvement is worth the higher token cost.
Use retrieval, caching, chunking, and selective context when they preserve quality at lower cost.
Input context
1,048,576
Max output
65,536
Requests above 200K input tokens remain within the context window but move into Gemini 2.5 Pro’s higher pricing tier.
06 / Thinking budget
Gemini 2.5 Pro Thinking Cannot Be Fully Disabled
Gemini 2.5 Pro is an always-thinking model: developers can reduce or increase the reasoning budget, but they cannot set the model to a true no-thinking mode.
128 to 32,768 Thinking Tokens
Google documents the following thinkingBudget behavior for Gemini 2.5 Pro:
- fixed range:
128to32768; thinkingBudget = -1: dynamic thinking;- dynamic thinking is the default;
thinkingBudget = 0: not supported.
The budget guides the maximum amount of reasoning the model can use.
It does not guarantee that the full budget will be consumed on every request.
This behavior is fundamentally different from Gemini 2.5 Flash, which can disable thinking completely.
For Pro workloads, the optimization question is therefore not “thinking or no thinking?”
It is:
How little reasoning can this task use without losing the advantage that justified choosing Pro?
Replay representative prompts at different budgets and measure:
- accepted-result quality;
- thinking tokens;
- total output tokens;
- latency;
- tool behavior;
- retries;
- cost.
Cost-sensitive Pro task
Use a smaller fixed budget when Pro-level capability is useful but deep reasoning does not improve the workload enough to justify extra latency.
Hard reasoning task
Use a larger budget or dynamic thinking for difficult coding, mathematics, STEM, and long-context synthesis.
07 / Multimodal analysis
Text, Images, Video, Audio, and PDFs in One Reasoning Context
Gemini 2.5 Pro can apply the same reasoning model across text and multimodal evidence instead of requiring separate specialist models for every input type.
Multimodal Reasoning Matters When Evidence Crosses Formats
A difficult task may combine:
- a written specification;
- screenshots;
- source code;
- diagrams;
- a PDF;
- recorded audio;
- video;
- external web evidence.
Gemini 2.5 Pro can receive these modalities together and produce a text result.
This is useful for:
- UI debugging from screenshots and code;
- technical-document analysis;
- video understanding;
- research across mixed sources;
- diagram interpretation;
- product or design review;
- multimodal question answering.
Google highlighted Gemini 2.5 Pro’s long-context and video-understanding performance during the model’s launch cycle.
For evaluation, test the entire multimodal workload rather than a single image prompt.
Measure whether the model correctly connects evidence across modalities and whether large media inputs trigger the >200K pricing tier.
- 01
Read
Analyze text, code, PDFs, and long documents inside the same working context.
- 02
See
Interpret images, screenshots, diagrams, and video alongside textual evidence.
- 03
Listen
Use audio input when spoken material is part of the analysis.
- 04
Synthesize
Reason across modalities and produce one text result grounded in the complete evidence set.
08 / Tools
Search, Maps, Files, Code, URLs, Functions, and Structured Outputs
Gemini 2.5 Pro combines deep reasoning with a mature tool surface for grounded research, computation, retrieval, and application actions.
Reasoning Can Be Connected to External Evidence
Google lists support for:
- Google Search grounding;
- Google Maps grounding;
- File Search;
- code execution;
- function calling;
- structured outputs;
- URL context;
- context caching;
- Batch inference;
- Flex inference;
- Priority inference.
This enables workflows such as:
- Search → compare evidence → structured answer;
- File Search → synthesize long documents;
- URL context → analyze external sources;
- code execution → calculate and verify;
- function calling → act through an application API.
Gemini 2.5 Pro does not support Gemini Live API on the base model ID.
Audio generation and image generation are also separate model capabilities rather than outputs of gemini-2.5-pro.
- Supported
Google Search
Ground reasoning in current web information.
- Supported
Google Maps
Use Maps grounding for supported location-aware workflows.
- Supported
File Search
Retrieve relevant content from indexed files.
- Supported
Code execution
Run code for calculations, transformations, and verification.
- Supported
Function calling
Invoke application-defined tools and actions.
- Supported
Structured outputs
Return schema-constrained machine-readable results.
- Supported
URL context
Read and reason over supplied URLs.
- Not listed
Live API
The base gemini-2.5-pro endpoint does not support Gemini Live API.
09 / Tuning
Gemini 2.5 Pro Supports Supervised Fine-Tuning on Vertex AI
Google Cloud supports supervised fine-tuning for Gemini 2.5 Pro, giving specialized applications a way to adapt the premium reasoning model to domain-specific behavior.
Tuning Should Solve a Specific Repeated Problem
Potential tuning workloads include:
- domain-specific coding;
- specialized terminology;
- structured domain responses;
- organization-specific output conventions;
- classification;
- entity extraction;
- specialized query generation.
For a Pro model, tuning economics deserve extra scrutiny.
The base inference price is already substantially higher than Flash tiers.
A tuned Pro model is most compelling when customization improves a high-value task that still requires Pro-level reasoning.
Before tuning, establish a prompt-only baseline.
Measure:
- accepted-result quality;
- prompt length;
- reasoning-token use;
- latency;
- training and maintenance overhead;
- regression risk;
- cost per successful task.
- 01
Baseline
Evaluate the untuned Pro model on a representative set of difficult production tasks.
- 02
Tune
Use supervised tuning when repeated specialized behavior is difficult to achieve reliably through prompting.
- 03
Compare
Measure whether tuning improves task completion enough to justify operational complexity and Pro-tier inference cost.
- 04
Monitor
Maintain regression tests across reasoning, formatting, coding, and domain-specific task classes.
10 / Lifecycle
Gemini 2.5 Pro Is Stable and Not Deprecated, but Access Is Restricted
Google still serves the stable Gemini 2.5 Pro API, but the 2.5 family is now oriented toward existing users rather than new projects.
Existing Workloads and New Projects Have Different Guidance
Google’s current Gemini API deprecation documentation states that:
gemini-2.5-prois not deprecated;- no API shutdown date has been announced;
- the stable API will continue to be served until further notice;
- access to Gemini 2.5 models is being limited to users who have actively used the family in the past.
For new projects, Google recommends current Gemini 3 models instead.
This makes Gemini 2.5 Pro a legacy production baseline, not an obsolete endpoint.
Existing systems may reasonably keep it when:
- its behavior is already validated;
- regression risk is expensive;
- tuning work has been invested;
- a migration has not yet demonstrated better task economics.
New systems should generally evaluate the current Gemini generation first.
- 01
Still stable
The Gemini API currently lists gemini-2.5-pro as a stable model.
- 02
Not deprecated
Google has announced no Gemini API shutdown date for the stable endpoint.
- 03
Access is restricted
Gemini 2.5 access is increasingly limited to users who have actively used the family before.
- 04
New projects should benchmark Gemini 3
Google recommends its latest model generation for new applications rather than starting fresh on Gemini 2.5.
11 / Migration
Migrating From Gemini 2.5 Pro to Newer Gemini Models
Moving away from Gemini 2.5 Pro requires more than changing the model ID because reasoning controls, pricing, agent behavior, and supported workflows have evolved in Gemini 3.
Re-Baseline the Workload Instead of Assuming Equivalence
Gemini 2.5 Pro uses thinkingBudget.
Its fixed range is 128 to 32,768 tokens and dynamic thinking is the default.
Gemini 3 models primarily use thinking_level.
That means a 2.5 Pro reasoning budget should not be mechanically translated into low, medium, or high.
Replay a representative workload and compare:
- first-pass quality;
- coding correctness;
- long-context retrieval;
- multimodal reasoning;
- tool-call behavior;
- reasoning-token usage;
- latency;
- total cost.
For difficult workloads, newer Pro-generation models are the natural capability comparison.
For tasks that do not consistently need premium reasoning, newer Flash models may offer better economics.
The migration goal should not be “find a model with the same name tier.”
It should be “find the least expensive current model that still meets the task’s acceptance criteria.”
- 01
Map reasoning behavior
Translate explicit 2.5 thinking budgets into empirically tested Gemini 3 reasoning levels.
- 02
Replay hard tasks
Use coding, STEM, long-context, multimodal, and tool-use cases that originally justified choosing Pro.
- 03
Challenge the tier
Test whether a newer Flash model now meets the same acceptance threshold at materially lower cost.
- 04
Measure completed-task economics
Include reasoning, retries, human correction, tool calls, and context pricing rather than comparing headline token rates.
12 / Evaluation
Gemini 2.5 Pro Strengths and Limitations
Gemini 2.5 Pro remains a strong historical and production baseline for deep reasoning, coding, long context, and multimodal analysis, but its always-on reasoning and premium pricing make workload selection critical.
Strengths
Deep reasoning
Designed specifically for difficult coding, mathematics, STEM, and multi-step analytical problems.
Strong coding capability
Google positioned 2.5 Pro as its leading Gemini 2.5 coding model for web apps, code transformation, debugging, and complex engineering.
1M-token multimodal context
Large codebases, documents, images, video, audio, retrieval, and tool results can share one reasoning context.
Rich production tool stack
Search, Maps, File Search, code execution, functions, structured outputs, URLs, caching, and multiple serving tiers support complex workflows.
What to consider
Thinking cannot be disabled
Gemini 2.5 Pro always reasons; the minimum fixed thinking budget is 128 tokens.
Premium token pricing
Standard output starts at $10 per million tokens, substantially above Flash and Flash-Lite models.
Higher rates above 200K input
Long prompts above 200K tokens increase Standard input pricing to $2.50 and output pricing to $15 per million tokens.
Legacy-access model
Google is restricting Gemini 2.5 access for new users and recommends current Gemini models for new projects.
Evaluate the premium reasoning baseline
Test Gemini 2.5 Pro on your hardest prompts
Replay coding, mathematics, STEM, long-document, repository, multimodal, tool-use, and research tasks to compare accepted-result quality, thinking-token usage, latency, context cost, and total cost per completed task.
Start FreeGemini 2.5 Pro is most useful when deeper reasoning materially improves completion. Benchmark it against lower-cost Flash models and newer Gemini generations before routing routine traffic to the Pro tier.
Common Questions
What is Gemini 2.5 Pro?
Gemini 2.5 Pro is Google’s premium Gemini 2.5 thinking model for complex coding, mathematics, STEM reasoning, large datasets, codebases, long documents, and multimodal analysis.
What is the Gemini 2.5 Pro model ID?
The stable model ID is gemini-2.5-pro.
When was Gemini 2.5 Pro released?
Google introduced Gemini 2.5 Pro in March 2025 and released the stable generally available gemini-2.5-pro endpoint on June 17, 2025.
Is Gemini 2.5 Pro stable?
Yes. The Gemini API currently lists gemini-2.5-pro as a stable model.
Is Gemini 2.5 Pro deprecated?
No. Google says the stable Gemini 2.5 models are not deprecated and will continue to be served until further notice through the API. No shutdown date is currently announced for gemini-2.5-pro.
Can new users access Gemini 2.5 Pro?
Google is limiting access to Gemini 2.5 models to users who have actively used them in the past and recommends newer Gemini models for new projects.
What is the Gemini 2.5 Pro context window?
Gemini 2.5 Pro supports up to 1,048,576 input tokens and up to 65,536 output tokens.
What input types does Gemini 2.5 Pro support?
Gemini 2.5 Pro accepts audio, images, video, text, and PDF input and generates text output.
What is the Gemini 2.5 Pro knowledge cutoff?
Google documents January 2025 as the knowledge cutoff.
How much does Gemini 2.5 Pro cost?
For Standard paid inference with prompts up to 200K tokens, pricing is $1.25 per 1M input tokens and $10 per 1M output tokens including thinking. Above 200K input tokens, pricing increases to $2.50 input and $15 output.
How much does Gemini 2.5 Pro context caching cost?
Cache reads cost $0.125 per 1M tokens for prompts up to 200K and $0.25 per 1M tokens above 200K. Cache storage is $4.50 per 1M tokens per hour.
What are Gemini 2.5 Pro Batch prices?
Batch pricing is $0.625 input and $5 output per million tokens for prompts up to 200K, and $1.25 input and $7.50 output above 200K.
What are Gemini 2.5 Pro Flex prices?
Flex uses the same core token rates as Batch: $0.625 input and $5 output up to 200K tokens, then $1.25 input and $7.50 output above 200K.
Does Gemini 2.5 Pro have a free tier?
Google currently lists free Standard token usage for Gemini 2.5 Pro in the Gemini Developer API, subject to provider limits and terms.
Does Gemini 2.5 Pro support thinking?
Yes. Gemini 2.5 Pro is an always-thinking model and uses dynamic thinking by default.
What is the Gemini 2.5 Pro thinking budget?
The supported fixed thinkingBudget range is 128 to 32,768 tokens. Setting thinkingBudget to -1 enables dynamic thinking, which is the default.
Can I disable thinking on Gemini 2.5 Pro?
No. Google explicitly documents that thinking cannot be disabled on Gemini 2.5 Pro. The minimum fixed thinking budget is 128 tokens.
Is Gemini 2.5 Pro good for coding?
Yes. Google positions Gemini 2.5 Pro as a leading coding model for complex code tasks, web development, code transformation, debugging, and agentic engineering workflows.
Is Gemini 2.5 Pro good for long documents and codebases?
Yes. Google specifically describes it as capable of analyzing large datasets, codebases, and documents using its 1,048,576-token context window.
Does Gemini 2.5 Pro support File Search?
Yes. Google’s current model specification lists File Search as supported.
What tools does Gemini 2.5 Pro support?
Google lists Search grounding, Maps grounding, File Search, code execution, function calling, structured outputs, URL context, and context caching.
Does Gemini 2.5 Pro support Live API?
No. The base gemini-2.5-pro endpoint does not support Gemini Live API.
Does Gemini 2.5 Pro support image generation or audio generation?
No. The base Gemini 2.5 Pro model accepts multimodal input but produces text. Google provides separate model endpoints for image and audio generation.
Does Gemini 2.5 Pro support tuning?
Yes on Vertex AI. Google Cloud lists Gemini 2.5 Pro among the models that support supervised fine-tuning.
How is Gemini 2.5 Pro different from Gemini 2.5 Flash?
Gemini 2.5 Pro targets deeper reasoning and stronger coding at substantially higher cost. Its thinking cannot be disabled and its fixed budget can reach 32,768 tokens. Gemini 2.5 Flash is cheaper, supports thinking budgets from 0 to 24,576 tokens, and can run with thinking disabled.
Why does Gemini 2.5 Pro cost more above 200K tokens?
Google uses a separate long-context pricing tier for prompts above 200K input tokens. Standard rates increase from $1.25/$10 to $2.50/$15 per million input/output tokens.
Should I use Gemini 2.5 Pro for a new application?
For a new application, Google recommends newer Gemini models. Gemini 2.5 Pro remains most relevant for established workloads that already depend on its validated deep-reasoning, coding, long-context behavior, or tuning investments.
Model information
Last updated
Specifications, pricing thresholds, context limits, thinking-budget behavior, coding and reasoning positioning, multimodal inputs, tool support, tuning, release status, access restrictions, and migration guidance on this page are based on official Google Gemini API, Google Cloud, and Google product documentation.