OpenAI flagship model
GPT-5.5
A flagship OpenAI model for complex professional work, coding, tool-heavy agents, grounded assistants, and long-context production workflows.
- Context window
- 1.05M
- tokens
- Max output
- 128K
- tokens
- Input
- $5.00
- per 1M tokens
- Cached input
- $0.50
- per 1M tokens
- Output
- $30.00
- per 1M tokens
01 / Overview
What GPT-5.5 Is
GPT-5.5 is an OpenAI flagship model for complex professional work, designed to handle demanding coding, agentic, retrieval, and customer-facing workflows with strong instruction following and polished output.
Built for Complex Production Work
OpenAI describes GPT-5.5 as a new class of intelligence for coding and professional work. Its guidance goes further by positioning the model for tool-heavy agents, grounded assistants, long-context retrieval, product-spec-to-plan workflows, and applications where execution quality matters as much as raw reasoning.
That combination makes GPT-5.5 especially relevant for production systems that need more than a single good answer. A workflow may need to interpret detailed requirements, retrieve supporting information, call tools, transform source material, and still produce a clear result that fits a product contract.
GPT-5.5 should be treated as its own model family rather than as a drop-in replacement for an older prompt stack. OpenAI recommends establishing a fresh baseline and then tuning reasoning effort, verbosity, tool descriptions, and output format against representative tasks.
- Flagship model for complex professional work.
- Strong fit for coding and tool-heavy agent workflows.
- Designed for grounded assistants and long-context retrieval.
- Supports a 1.05M-token context window.
- Highly steerable for structured and polished outputs.
- Provider
- OpenAI
- Family
- GPT-5.5
- Positioning
- Complex professional work
- Knowledge cutoff
- Dec 1, 2025
- Input modalities
- Text, Image
- Output modality
- Text
- Default reasoning
- Medium
02 / Use cases
Where GPT-5.5 Fits Best
GPT-5.5 is strongest when the application needs a model to combine reasoning, tools, retrieval, and high-quality output across a complete professional workflow rather than solve a narrow isolated prompt.
From Requirements to Finished Deliverables
One of the model's clearest use cases is turning an ambiguous or detailed specification into an executable plan. Product teams can use it to interpret requirements, identify dependencies, produce implementation steps, and refine the result against additional constraints.
Coding is another major workload. GPT-5.5 can participate in software-engineering tasks that involve reading code, using shell or patching tools, understanding broader project context, and producing changes that fit an existing system.
Grounded assistants and research-style workflows also benefit from its large context window and tool support. Instead of relying only on the model's internal knowledge, an application can retrieve files or external information, keep the relevant evidence in context, and ask the model to synthesize a response grounded in that material.
- Complex coding and software-engineering workflows.
- Tool-heavy agents that need several actions to complete a task.
- Grounded assistants that work from retrieved evidence.
- Product-spec-to-plan workflows and coordinated deliverables.
- Customer-facing experiences where response polish and instruction following matter.
- 01
Coding and engineering
Work across code, requirements, tools, and implementation constraints in substantial technical tasks.
- 02
Tool-heavy agents
Coordinate search, files, code execution, shell, patching, computer use, and application-defined tools.
- 03
Grounded assistants
Answer from retrieved documents and evidence instead of relying only on model knowledge.
- 04
Spec-to-plan workflows
Turn product requirements, source material, and constraints into structured plans and polished deliverables.
03 / Pricing
GPT-5.5 Pricing
GPT-5.5 standard pricing is $5.00 per million input tokens, $0.50 per million cached input tokens, and $30.00 per million output tokens.
Price the Complete Workflow, Not Just the Prompt
For professional workloads, total inference cost is often shaped by much more than the initial user message. Large system instructions, retrieved files, conversation history, tool definitions, and long generated deliverables can all contribute to token usage.
Output is the most expensive standard token category for GPT-5.5, so response length deserves explicit control. OpenAI recommends treating final-answer length separately from reasoning quality and using verbosity controls or concrete output requirements when a product needs predictable response size.
Prompt caching can reduce the cost of repeated stable context. GPT-5.5 uses its earlier-model caching behavior: eligible long prefixes can be reused at the cached-input rate, and OpenAI does not charge an additional cache-write rate for this model.
- $5.00 per 1M standard input tokens.
- $0.50 per 1M cached input tokens.
- $30.00 per 1M output tokens.
- No additional prompt-cache write charge for GPT-5.5.
- Tool-specific operations may have separate fees.
1M tokens · USD
- Input
- $5.00
- Cached input
- $0.50
- Output
- $30.00
Example: 10K input + 2K output
- Input cost
- $0.0500
- Output cost
- $0.0600
- Estimated total
- $0.1100
04 / Context
A 1.05M-Token Context Window for Retrieval and Complex Projects
GPT-5.5 supports a 1,050,000-token context window and up to 128,000 output tokens, giving applications room for large document sets, code, conversation state, and retrieved evidence.
Long Context Works Best When It Is Grounded and Deliberate
A large window is particularly useful for grounded assistants and retrieval-heavy applications. The model can work with substantial source material while preserving the instructions and context needed to produce a coherent answer.
For coding workflows, that may include source files, architecture notes, issue history, logs, and implementation requirements. For professional research, it may include reports, policies, customer records, or several retrieved documents that need to be reconciled.
The window should still be treated as a budget rather than a target. More context increases cost and can add irrelevant information. Above 272K input tokens, GPT-5.5 moves into higher long-context pricing, making retrieval quality and context selection financially important as well as technically important.
GPT-5.5 also supports extended prompt caching. Keeping reusable instructions and reference material stable at the beginning of a request can reduce input cost on repeated workflows.
- 1,050,000-token context window.
- Maximum output of 128,000 tokens.
- Well suited to long-context retrieval and grounded assistants.
- Extended prompt caching can reduce repeated context cost.
- Requests above 272K input tokens use higher pricing.
Context window
1,050,000
Max output
128,000
A production context may include system instructions, retrieved files, source documents, code, tool definitions, conversation history, and the current task.
05 / Reasoning
Reasoning Effort from None to XHigh
GPT-5.5 supports none, low, medium, high, and xhigh reasoning effort, with medium as the default.
Tune Reasoning and Output Separately
Reasoning effort controls how much inference effort the model applies to a task, while verbosity controls how much of the final result is presented to the user. OpenAI explicitly recommends treating these as separate concerns.
That distinction is useful in production. A difficult analysis may benefit from higher reasoning but still require a concise final answer. A customer-facing explanation may need polished formatting and more detail without necessarily requiring maximum reasoning.
For GPT-5.5, the best setting should be established against representative tasks. Start with a baseline, evaluate quality and cost, then increase or decrease reasoning effort only where the data supports the change.
When reasoning effort is not none, some sampling parameters such as temperature and top_p are not supported, so reasoning configuration also affects how the API request should be constructed.
Structured production task
Test none or low reasoning for well-defined tasks where the required output and decision process are straightforward.
Complex professional task
Use higher reasoning when coding, planning, synthesis, or tool-heavy execution involves multiple interacting constraints.
06 / Capabilities
Capabilities for Tool-Heavy Production Systems
GPT-5.5 supports text and image input, streaming, function calling, structured outputs, prompt caching, and a broad Responses API toolset for coding and agentic workflows.
Built to Work with Tools and Grounded Context
OpenAI lists web search, file search, image generation, code interpreter, hosted shell, Apply Patch, Skills, computer use, MCP, and Tool Search as supported Responses API tools for GPT-5.5.
This makes the model suitable for systems where the answer should be grounded in files or external information, or where the model needs to take actions instead of only describing them. A coding agent can inspect and patch code. A research assistant can retrieve evidence. A workflow can call application-defined functions and return structured data that downstream software can validate.
Structured Outputs are particularly useful when a product requires a stable schema. OpenAI recommends using the API's structured-output capability instead of describing the expected schema only in the prompt.
The model accepts text and image input and produces text output. Direct audio and video modalities are not supported. Fine-tuning is currently unavailable.
- Text input and output with image input.
- Streaming responses.
- Function calling and structured outputs.
- Web search and file search.
- Image generation and code interpreter.
- Hosted shell and Apply Patch.
- Skills, computer use, MCP, and Tool Search.
- Fine-tuning is currently not supported.
- Supported
Text and vision
Accept text and image input and generate text output.
- Supported
Structured integration
Use function calling and Structured Outputs for predictable application behavior.
- Supported
Grounded retrieval
Use web search and file search to work from external evidence and source material.
- Supported
Coding tools
Use code interpreter, hosted shell, and Apply Patch in technical workflows.
- Supported
Computer use
Interact with graphical environments as part of supported Responses API workflows.
- Supported
Extensible agents
Use Skills, MCP, Tool Search, image generation, and other supported tools.
07 / Evaluation
Strengths and Limitations
GPT-5.5 is a strong fit for complex professional workflows, but its higher token cost and earlier-generation caching behavior make workload-specific evaluation important before production deployment.
Strengths
Complex professional work
OpenAI positions GPT-5.5 for coding, agents, grounded assistants, long-context retrieval, and polished customer-facing workflows.
Large context window
A 1.05M-token context supports substantial documents, code, retrieved evidence, and project state.
Broad tool support
Search, files, code execution, shell, patching, computer use, MCP, Skills, and Tool Search support end-to-end workflows.
Strong output steerability
Reasoning effort, verbosity, Structured Outputs, and explicit format requirements can be tuned independently for product needs.
What to consider
Premium token pricing
At $5 input and $30 output per million tokens, simpler or high-volume workloads may be more economical on newer lower-cost model tiers.
Long-context pricing multiplier
Prompts above 272K input tokens are charged at 2× input and 1.5× output pricing for the full session.
Earlier prompt-caching model
GPT-5.5 uses extended automatic prompt caching rather than the explicit cache-breakpoint controls introduced in GPT-5.6 and later.
Fine-tuning is unavailable
The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.5.
Evaluate complex workflows
Test GPT-5.5 with your real production tasks
Run representative coding, retrieval, agentic, and professional workflows with your actual prompts and context, then inspect quality, token usage, cost, and context consumption in EidoStack.
Start FreeConnect your own OpenAI API key and evaluate GPT-5.5 under the same reasoning, prompt, context, and tool conditions your application will use.
Common Questions
What is GPT-5.5?
GPT-5.5 is an OpenAI flagship model for complex professional work. It is designed for coding, tool-heavy agents, grounded assistants, long-context retrieval, product-spec-to-plan workflows, and customer-facing tasks where execution quality matters.
How much does GPT-5.5 cost?
Standard pricing is $5.00 per 1M input tokens, $0.50 per 1M cached input tokens, and $30.00 per 1M output tokens. Tool-specific charges may apply separately.
What is the context window of GPT-5.5?
GPT-5.5 has a 1,050,000-token context window and supports up to 128,000 output tokens.
What reasoning levels does GPT-5.5 support?
GPT-5.5 supports none, low, medium, high, and xhigh reasoning effort. Medium is the default.
What is the knowledge cutoff for GPT-5.5?
OpenAI lists December 1, 2025 as the knowledge cutoff for GPT-5.5.
Does GPT-5.5 support image input?
Yes. GPT-5.5 supports text and image input and produces text output. Direct audio and video modalities are not supported.
What tools does GPT-5.5 support?
OpenAI currently lists web search, file search, image generation, code interpreter, hosted shell, Apply Patch, Skills, computer use, MCP, and Tool Search as supported Responses API tools.
Does GPT-5.5 support prompt caching?
Yes. GPT-5.5 supports extended prompt caching. Eligible repeated prefixes can be billed at the cached-input rate, and there is no additional cache-write charge for this model.
What workloads are a good fit for GPT-5.5?
GPT-5.5 is well suited to complex coding, tool-heavy agents, grounded assistants, long-context retrieval, professional planning, research synthesis, and customer-facing workflows that need polished, controlled output.
Can GPT-5.5 be fine-tuned?
No. The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.5.
Model information
Last updated
The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.5. Provider pricing, limits, processing options, prompt-caching behavior, supported tools, and availability may change, so production assumptions should be checked against the latest provider documentation.