OpenAI Pro reasoning model
GPT-5 Pro
A higher-compute GPT-5 variant built to think harder and produce more precise answers on difficult problems, with high reasoning effort and long-running Responses API execution.
- Context window
- 400K
- tokens
- Max output
- 272K
- tokens
- Input
- $15.00
- per 1M tokens
- Output
- $120.00
- per 1M tokens
- Reasoning
- High
- only
01 / Overview
What GPT-5 Pro Is
GPT-5 Pro is the higher-compute version of GPT-5, designed to spend more inference effort on difficult problems and deliver consistently stronger, more precise answers when quality matters more than speed or token cost.
A Precision-First GPT-5 Variant
GPT-5 Pro was released on October 6, 2025 as a specialized reasoning tier above the standard GPT-5 model.
Its purpose is straightforward: use more compute to think harder.
That makes GPT-5 Pro structurally different from a general-purpose model where developers can freely move between fast and deep reasoning settings. GPT-5 Pro only supports reasoning.effort: high. There is no low-cost, low-reasoning operating mode to tune for routine requests.
OpenAI also designed the model for the Responses API rather than ordinary low-latency request patterns. Some difficult requests can take several minutes, and OpenAI recommends background mode to avoid timeouts.
This creates a clear workload boundary. GPT-5 Pro is most appropriate when the value of a better answer justifies higher cost and longer execution time: hard analysis, difficult reasoning, high-value decisions, and escalation paths where a cheaper model has already failed.
The model is now deprecated, so its strongest current role is as a migration and regression baseline for existing Pro workloads.
- Higher-compute GPT-5 reasoning variant.
- High reasoning effort is mandatory.
- Responses API only.
- Some requests can take several minutes.
- Background mode is recommended for long-running work.
- Deprecated with a defined migration path.
- Provider
- OpenAI
- Family
- GPT-5 Pro
- Positioning
- Precision-first reasoning
- Status
- Deprecated
- Knowledge cutoff
- Sep 30, 2024
- Input modalities
- Text, Image
- Output modality
- Text
02 / Use cases
Where GPT-5 Pro Fit Best
GPT-5 Pro is best understood as an escalation model: use it for difficult, high-value work where additional reasoning quality can justify substantially higher inference cost and longer response time.
Escalate Hard Problems Instead of Routing Everything to Pro
The economics of GPT-5 Pro favor selective use.
At $15 per million input tokens and $120 per million output tokens, it is far more expensive than the standard GPT-5 tier. It also cannot reduce reasoning effort below high.
That combination makes routine classification, extraction, rewriting, routing, and simple question answering poor fits. Those tasks generally do not benefit enough from Pro-level inference to justify the premium.
The stronger pattern is escalation. A lower-cost model handles normal traffic, while GPT-5 Pro receives the small subset of cases that are difficult, ambiguous, consequential, or repeatedly unsuccessful.
Examples include complex technical investigations, difficult planning, synthesis across a large evidence set, high-value analytical work, and challenging decisions where the cost of an incorrect or shallow answer is greater than the additional inference expense.
Because the model can run for several minutes, asynchronous or background workflows are a more natural fit than latency-sensitive interactive UI paths.
- 01
Difficult investigations
Escalate technically complex or ambiguous problems that require sustained reasoning across substantial evidence.
- 02
High-value analysis
Use additional inference compute when a more precise answer can materially improve an important decision.
- 03
Complex synthesis
Reason across long documents, constraints, competing evidence, and multiple intermediate conclusions.
- 04
Escalation tier
Route only the hardest cases to Pro after a faster or cheaper model cannot meet the required quality threshold.
03 / Pricing
GPT-5 Pro Pricing
GPT-5 Pro is priced at $15.00 per million input tokens and $120.00 per million output tokens, reflecting its position as a high-compute reasoning model rather than a general-purpose default.
Optimize for Value per Solved Problem, Not Cheapest Token
GPT-5 Pro's price changes the evaluation question.
For a routine workload, the relevant question is often how to minimize cost while maintaining acceptable quality. For a Pro workload, the question is whether extra reasoning materially increases the probability of solving an expensive problem correctly.
Output is especially significant at $120 per million tokens. Long reasoning-heavy tasks can therefore accumulate meaningful cost even when the visible final answer is concise.
OpenAI's current GPT-5 Pro model card does not list a separate cached-input price. Cost models should not assume the discounted cached-input rate used by standard GPT-5.
A practical production architecture can use a cheaper model for first-pass processing and reserve GPT-5 Pro for explicit escalation conditions. This limits premium inference to cases where it has the highest expected value.
- $15.00 per 1M input tokens.
- $120.00 per 1M output tokens.
- No separate cached-input rate is listed on the current model card.
- High reasoning is always enabled.
- Measure cost per successfully resolved hard task.
1M tokens · USD
- Input
- $15.00
- Output
- $120.00
Example: 10K input + 2K output
- Input cost
- $0.1500
- Output cost
- $0.2400
- Estimated total
- $0.3900
04 / Context
400K Context with Up to 272K Output
GPT-5 Pro provides a 400,000-token context window and an unusually large 272,000-token maximum output, giving high-compute reasoning substantial room for both source material and generated work.
Large Output Capacity Changes the Shape of Pro Workloads
The 400K context window can accommodate substantial source material: technical documentation, specifications, code, retrieved evidence, prior reasoning state, and detailed task instructions.
More distinctive is the 272K maximum output limit.
That ceiling is much larger than the 128K output limit used by many later GPT models. It gives GPT-5 Pro room for very long generated artifacts or reasoning-heavy workflows when the application requires them.
A high maximum does not mean applications should routinely request extremely long responses. At Pro output pricing, unnecessary generation is expensive, and additional context can also introduce irrelevant information.
For evaluation, measure how much context the model actually needs and how much output the successful workflow actually consumes. Long-context capacity should support hard tasks rather than become a reason to send unfiltered data.
Context window
400,000
Max output
272,000
GPT-5 Pro combines a 400K context window with a 272K maximum output. For expensive Pro workloads, curate context and control generated length rather than treating both limits as targets.
05 / Reasoning
High Reasoning Is the Only Operating Mode
GPT-5 Pro defaults to and only supports reasoning.effort: high, making deep inference a fixed property of the model rather than a per-request performance dial.
Pro Removes the Fast-vs-Deep Reasoning Choice
Standard reasoning models often expose several effort levels so developers can balance quality, latency, and cost.
GPT-5 Pro deliberately removes that choice.
Every request uses high reasoning effort. This simplifies the model's role: it is intended for work that already justifies substantial inference rather than workloads that need dynamic routing between quick and deep responses.
The tradeoff is latency. OpenAI notes that some GPT-5 Pro requests can take several minutes. Applications should therefore be designed around long-running execution instead of assuming an immediate synchronous response.
Background mode is particularly important for these cases because it allows a difficult request to continue without depending on a single long-lived client connection.
For migration, preserve the semantic role of Pro reasoning. OpenAI recommends gpt-5.6-sol with reasoning.mode: pro for the deprecated dated snapshot rather than treating ordinary reasoning settings as automatically equivalent.
Hard analytical task
Use GPT-5 Pro when deeper inference is a requirement of the workload rather than an optional optimization.
Long-running execution
Plan for requests that may take several minutes and use Responses API background mode when appropriate.
06 / Capabilities
Responses API, Structured Output, and Long-Running Work
GPT-5 Pro supports streaming, function calling, structured outputs, text and image input, and is designed specifically around the Responses API for advanced multi-turn reasoning behavior.
API Design Reflects the Model's Long-Running Nature
OpenAI states that GPT-5 Pro is available in the Responses API only.
This is important for implementation. A migration from a conventional chat request should not assume that GPT-5 Pro behaves like a drop-in low-latency Chat Completions model.
The model supports function calling and structured outputs, so applications can still integrate Pro reasoning into deterministic workflows and machine-readable pipelines.
Streaming is also supported, although the potentially long execution time means streaming alone does not remove timeout concerns. Background mode is the stronger pattern for requests that may run for minutes.
Text and image are supported as inputs, while output is text.
One explicit limitation is Code Interpreter: OpenAI states that GPT-5 Pro does not support it. Fine-tuning and predicted outputs are also listed as unsupported.
- Supported
Responses API
GPT-5 Pro is designed for the Responses API and its advanced multi-turn model interaction patterns.
- Supported
Background mode
Recommended for difficult requests that may require several minutes to finish.
- Supported
Function calling
Connect high-compute reasoning to application-defined tools and actions.
- Supported
Structured outputs
Return schema-constrained results for downstream application logic.
- Not listed
Code Interpreter
OpenAI explicitly states that GPT-5 Pro does not support Code Interpreter.
- Not listed
Fine-tuning
Fine-tuning is not supported for GPT-5 Pro.
07 / Evaluation
Strengths and Limitations
GPT-5 Pro prioritizes reasoning quality over flexibility, latency, and cost: it can be valuable for genuinely difficult problems, but high-only reasoning, premium pricing, long execution times, and deprecation make it unsuitable as a default model for ordinary traffic.
Where GPT-5 Pro stands out
Higher-compute reasoning
The model is explicitly designed to think harder and provide more precise answers on difficult problems.
Clear escalation role
High-only reasoning makes GPT-5 Pro easy to position as a premium fallback for cases where standard models fail quality thresholds.
Large output capacity
The 272K maximum output provides unusually large generation headroom for complex long-form workflows.
Structured application integration
Function calling and structured outputs allow high-compute reasoning to participate in machine-readable production workflows.
What to consider
Deprecated
GPT-5 Pro is deprecated, and the gpt-5-pro-2025-10-06 snapshot is scheduled for shutdown on December 11, 2026.
Premium token cost
$15 input and $120 output per million tokens make indiscriminate routing expensive.
High reasoning only
The model cannot be tuned down to a lower reasoning effort for faster or cheaper routine tasks.
Potentially long latency
OpenAI notes that difficult requests may take several minutes, requiring background-oriented application design.
No Code Interpreter
OpenAI explicitly lists Code Interpreter as unsupported for GPT-5 Pro.
Preserve your Pro baseline
Evaluate GPT-5 Pro before migrating
Replay difficult production tasks against GPT-5 Pro and its replacement, then compare answer quality, reasoning behavior, token usage, latency, and total cost in EidoStack.
Start FreeGPT-5 Pro is deprecated. OpenAI recommends gpt-5.6-sol with reasoning.mode: pro for the gpt-5-pro-2025-10-06 migration path.
Common Questions
What is GPT-5 Pro?
GPT-5 Pro is a higher-compute version of GPT-5 designed to think harder and provide more precise answers on difficult problems.
Is GPT-5 Pro deprecated?
Yes. OpenAI currently marks GPT-5 Pro as deprecated.
When will GPT-5 Pro shut down?
OpenAI's deprecation schedule lists December 11, 2026 as the shutdown date for the gpt-5-pro-2025-10-06 snapshot.
What replaces GPT-5 Pro?
For gpt-5-pro-2025-10-06, OpenAI recommends gpt-5.6-sol with reasoning.mode: pro.
How much does GPT-5 Pro cost?
OpenAI lists GPT-5 Pro at $15.00 per 1M input tokens and $120.00 per 1M output tokens.
Does GPT-5 Pro have cached-input pricing?
The current GPT-5 Pro model card does not list a separate cached-input price, so cost estimates should not assume the cached-input discount available on standard GPT-5.
What is the GPT-5 Pro context window?
GPT-5 Pro has a 400,000-token context window.
What is the maximum output of GPT-5 Pro?
GPT-5 Pro supports up to 272,000 output tokens.
What reasoning effort does GPT-5 Pro support?
GPT-5 Pro defaults to and only supports high reasoning effort.
Why can GPT-5 Pro take several minutes?
GPT-5 Pro uses more compute to reason through difficult problems. OpenAI notes that some requests can take several minutes and recommends background mode to avoid timeouts.
Which API does GPT-5 Pro use?
OpenAI states that GPT-5 Pro is available in the Responses API only.
Does GPT-5 Pro support function calling?
Yes. The official model card lists function calling as supported.
Does GPT-5 Pro support structured outputs?
Yes. Structured outputs are supported.
Does GPT-5 Pro support Code Interpreter?
No. OpenAI explicitly states that GPT-5 Pro does not support Code Interpreter.
Can GPT-5 Pro be fine-tuned?
No. The official model card lists fine-tuning as unsupported.
When should an application use a Pro model?
A Pro model is most appropriate as an escalation tier for difficult, high-value tasks where improved reasoning quality can justify substantially higher cost and longer latency.
Model information
Last updated
The specifications, pricing, capabilities, API behavior, and lifecycle information on this page are based on OpenAI's official GPT-5 Pro model card, API changelog, and deprecation documentation.