OpenAI GPT-5 model
GPT-5.1
A GPT-5 generation model built around efficient adaptive reasoning, responsive coding, agentic tool use, and a no-reasoning mode for latency-sensitive tasks.
- Context window
- 400K
- tokens
- Max output
- 128K
- tokens
- Input
- $1.25
- per 1M tokens
- Cached input
- $0.125
- per 1M tokens
- Output
- $10.00
- per 1M tokens
01 / Overview
What GPT-5.1 Is
GPT-5.1 is an OpenAI model for coding and agentic tasks that introduced a more adaptive reasoning strategy: it can spend fewer tokens on straightforward work, think longer on harder problems, or disable reasoning entirely for latency-sensitive requests.
Efficient Reasoning Was the Core Design Goal
GPT-5.1 was released in November 2025 as the next step in the GPT-5 series, with a focus on balancing intelligence and speed across coding and agentic workloads.
Instead of spending a similar amount of reasoning effort on every request, GPT-5.1 was trained to adapt its thinking more dynamically to task complexity. Straightforward requests could complete with less hidden reasoning, while difficult tasks could remain persistent and explore more options before returning an answer.
The release also added a none reasoning mode. This made GPT-5.1 useful in architectures that wanted one model family to cover both fast tool-calling requests and deeper reasoning tasks.
OpenAI paired that reasoning design with stronger coding behavior, improved steerability, extended prompt caching, and new code-oriented tools such as Apply Patch and shell workflows.
GPT-5.1 is now deprecated. OpenAI announced deprecation on October 1, 2026, with API shutdown scheduled for April 1, 2027. GPT-6 Sol is the recommended replacement.
- Designed for coding and agentic tasks.
- Adaptive reasoning reduces unnecessary thinking on easier work.
- Supports a no-reasoning mode for fast requests.
- Introduced extended prompt caching up to 24 hours.
- Deprecated with GPT-6 Sol as the recommended replacement.
- Provider
- OpenAI
- Family
- GPT-5.1
- Positioning
- Coding & agentic
- Status
- Deprecated
- Knowledge cutoff
- Sep 30, 2024
- Input modalities
- Text, Image
- Default reasoning
- None
02 / Use cases
Where GPT-5.1 Fits Best
GPT-5.1 was strongest in systems that mixed fast routine actions with harder coding and agentic tasks, because one model could move from no-reasoning execution to deeper deliberation as workload complexity changed.
One Model Across Fast and Difficult Agent Steps
Many agent workflows contain a mixture of tasks.
One step may simply choose a tool, classify a request, or make a small code edit. Another may require debugging a failure across several files, planning a refactor, or reasoning through an ambiguous requirement.
GPT-5.1's reasoning controls were well suited to that pattern. Developers could use none or low for responsive operations, then allocate more inference effort to harder cases.
Coding was a major focus of the release. OpenAI described improvements in coding personality, steerability, code quality, front-end generation, and user-facing progress updates during sequences of tool calls.
The model also targeted tool-heavy agents. OpenAI specifically highlighted improved parallel tool calling in no-reasoning mode and introduced Apply Patch and shell workflows with GPT-5.1.
Because the model is deprecated, these use cases are now most useful as migration benchmarks for existing systems rather than reasons to start a new GPT-5.1 integration.
- Responsive coding assistants with frequent short turns.
- Agent workflows that mix simple and difficult steps.
- Parallel tool calling at low reasoning effort.
- Multi-file code editing and iterative repair.
- Legacy GPT-5.1 workload benchmarking before migration.
- 01
Fast tool steps
Use no-reasoning behavior for latency-sensitive actions, routing, and straightforward tool calls.
- 02
Coding iteration
Handle quick code edits, repository changes, frontend work, and interactive developer loops.
- 03
Harder engineering
Increase reasoning effort when debugging, planning, or multi-file changes require more persistence.
- 04
Agent orchestration
Combine function calls, shell execution, patches, search, and repeated model turns inside a controlled workflow.
03 / Pricing
GPT-5.1 Pricing
GPT-5.1 is priced at $1.25 per million input tokens, $0.125 per million cached input tokens, and $10.00 per million output tokens.
Reasoning Efficiency Was Part of the Cost Story
GPT-5.1's pricing matched GPT-5 at launch, but OpenAI positioned the newer model as more token-efficient because it could spend less time reasoning on easy tasks.
That distinction matters in agent systems. Two models with the same nominal token price can produce different end-to-end costs if one uses fewer reasoning tokens, finishes in fewer turns, or avoids unnecessary actions.
Cached input was priced at one tenth of standard input. GPT-5.1 also introduced extended prompt caching with retention of up to 24 hours, which made cached context more useful for long-running conversations, coding sessions, and retrieval workflows.
For migration analysis, preserve complete workflow measurements. Compare not just input and output prices, but also reasoning usage, number of turns, retries, tool calls, and completion rate.
- $1.25 per 1M input tokens.
- $0.125 per 1M cached input tokens.
- $10.00 per 1M output tokens.
- Extended prompt caching could retain eligible prefixes for up to 24 hours.
- Evaluate total cost per completed agent or coding task.
1M tokens · USD
- Input
- $1.25
- Cached input
- $0.125
- Output
- $10.00
Example: 10K input + 2K output
- Input cost
- $0.0125
- Output cost
- $0.0200
- Estimated total
- $0.0325
04 / Context
A 400K Context Window for Agentic Work
GPT-5.1 provides a 400,000-token context window and supports up to 128,000 output tokens, giving coding and agent systems substantial room for instructions, source files, tool definitions, conversation state, and intermediate results.
Stable Prefixes Could Remain Cached Longer
The context window itself was large enough for significant coding and professional workloads, but GPT-5.1's more distinctive context feature was extended prompt caching.
At launch, OpenAI allowed developers to request cache retention of up to 24 hours. That was designed to improve latency and cost for follow-up requests that shared the same prompt prefix.
Coding sessions are a clear example. System instructions, repository guidance, tool definitions, and stable project context may remain unchanged while the user and agent iterate over several steps.
The same pattern appears in support agents, research workflows, and knowledge assistants.
A large context window still benefits from deliberate selection. Sending every available file or conversation turn can increase token usage and dilute the relevant signal even when everything technically fits.
- 400,000-token context window.
- Maximum output of 128,000 tokens.
- Extended caching supported up to 24 hours at launch.
- Useful for repeated coding and agent interactions with stable prefixes.
- Context selection remains important for cost and relevance.
Context window
400,000
Max output
128,000
A GPT-5.1 coding session could reuse stable system instructions, tool schemas, repository guidance, and project context while new user requests and tool results changed between turns.
05 / Reasoning
Adaptive Reasoning from None to High
GPT-5.1 supports none, low, medium, and high reasoning effort, with none as the default.
No Reasoning Was a First-Class Operating Mode
The none setting was one of the defining additions in GPT-5.1.
OpenAI designed it for latency-sensitive tasks that did not require deep deliberation. In this mode, the model behaved more like a non-reasoning model while retaining GPT-5.1's coding, instruction-following, and tool-calling capabilities.
For more complex requests, developers could move to low, medium, or high.
OpenAI's launch guidance recommended low or medium for higher-complexity tasks and high when intelligence and reliability mattered more than speed.
The model also adapted its thinking within a chosen reasoning level. OpenAI described GPT-5.1 as spending fewer reasoning tokens on straightforward tasks while remaining persistent on harder tasks.
Latency-sensitive step
Use none for simple tool calls, quick edits, routing, and other tasks where extra deliberation does not improve the result.
Hard agent task
Increase to low, medium, or high when debugging, planning, research, or multi-step execution requires more reliability.
06 / Capabilities
Coding, Tool Calling, Patching, and Shell Workflows
GPT-5.1 supports streaming, function calling, structured outputs, text and image input, and was launched with new Apply Patch and shell workflows designed for iterative agentic coding.
A Model Designed to Do More Than Return Code
The official model card lists streaming, function calling, and structured outputs as supported.
GPT-5.1 also accepts image input, which can be useful when coding or agent tasks include screenshots, visual references, diagrams, or interface state.
At launch, OpenAI introduced Apply Patch with GPT-5.1. The tool allowed the model to create, update, and delete files using structured diffs that an application could apply and report back on.
OpenAI also introduced a shell workflow that allowed GPT-5.1 to propose commands, receive execution results, and continue through a plan-execute loop.
Web search was supported with no-reasoning mode as well, and the release emphasized stronger parallel tool calling.
Because GPT-5.1 is deprecated, current tool availability should be checked before relying on these historical integrations, and new systems should target GPT-6 Sol instead.
- Streaming supported.
- Function calling supported.
- Structured outputs supported.
- Text and image input with text output.
- Apply Patch introduced with GPT-5.1.
- Shell workflows introduced with GPT-5.1.
- Web search supported in documented GPT-5.1 workflows.
- Fine-tuning and predicted outputs are not supported.
- Supported
Function calling
Connect model decisions to application-defined actions and parallel tool workflows.
- Supported
Structured outputs
Return machine-readable responses that conform to an application schema.
- Supported
Apply Patch
Historically introduced with GPT-5.1 for structured create, update, and delete operations across code files.
- Supported
Shell workflow
Historically introduced with GPT-5.1 for command execution loops where the integration runs commands and returns results.
- Supported
Image input
Use screenshots and other visual context alongside text and code.
- Not listed
Fine-tuning
OpenAI lists fine-tuning as unsupported for GPT-5.1.
07 / Evaluation
Strengths and Limitations
GPT-5.1 represented an important shift toward more efficient adaptive reasoning for coding and agents, but its deprecated lifecycle now makes it primarily a migration and regression-testing target rather than a model for new production systems.
Historical strengths
Adaptive reasoning
GPT-5.1 was trained to spend less reasoning effort on easier tasks while remaining persistent when the task became difficult.
True no-reasoning mode
The none setting gave developers a low-latency operating mode inside the same model family used for deeper agentic work.
Coding-focused improvements
OpenAI emphasized better steerability, code quality, frontend generation, tool-call updates, and iterative developer experience.
Extended prompt caching
Up to 24-hour cache retention improved the economics of repeated multi-turn sessions with stable prompt prefixes.
What to consider
Deprecated lifecycle
GPT-5.1 was deprecated on October 1, 2026 and is scheduled to shut down on April 1, 2027. OpenAI recommends GPT-6 Sol.
Older knowledge cutoff
The official model card lists September 30, 2024 as the knowledge cutoff.
No xhigh reasoning
GPT-5.1 supports none, low, medium, and high, while later GPT generations extend the reasoning range further.
Newer agent models supersede it
Current OpenAI models provide newer capabilities, longer support horizons, and updated tool ecosystems for coding and agent workloads.
Preserve behavior before migration
Benchmark GPT-5.1 before moving your coding and agent workloads
Capture representative coding, tool-calling, low-latency, and long-context cases, then compare quality, reasoning behavior, token usage, and cost against GPT-6 Sol in EidoStack.
Start FreeGPT-5.1 is deprecated. OpenAI recommends migrating production workloads to GPT-6 Sol before the April 1, 2027 API shutdown.
Common Questions
What is GPT-5.1?
GPT-5.1 is an OpenAI model designed for coding and agentic tasks with configurable reasoning. It introduced more adaptive reasoning behavior, a no-reasoning mode, extended prompt caching, and new coding-oriented tool workflows.
Is GPT-5.1 deprecated?
Yes. OpenAI deprecated GPT-5.1 on October 1, 2026 and plans to remove it from the API on April 1, 2027.
What should replace GPT-5.1?
OpenAI lists GPT-6 Sol as the recommended replacement for GPT-5.1.
How much does GPT-5.1 cost?
OpenAI lists $1.25 per 1M input tokens, $0.125 per 1M cached input tokens, and $10.00 per 1M output tokens.
What is the context window of GPT-5.1?
GPT-5.1 has a 400,000-token context window and supports up to 128,000 output tokens.
What reasoning levels does GPT-5.1 support?
GPT-5.1 supports none, low, medium, and high reasoning effort. None is the default.
What is adaptive reasoning in GPT-5.1?
OpenAI trained GPT-5.1 to vary how much effort it spends thinking based on task complexity. Straightforward tasks can use fewer reasoning tokens, while harder tasks can remain more persistent.
What is GPT-5.1 no-reasoning mode?
Setting reasoning effort to none disables extended reasoning for latency-sensitive tasks while retaining GPT-5.1's general intelligence, coding, and tool-calling behavior.
What is the knowledge cutoff for GPT-5.1?
OpenAI lists September 30, 2024 as the knowledge cutoff for GPT-5.1.
Does GPT-5.1 support image input?
Yes. GPT-5.1 accepts text and image input and produces text output. Direct audio and video modalities are not supported.
Does GPT-5.1 support structured outputs?
Yes. The official model card lists structured outputs, function calling, and streaming as supported.
What was extended prompt caching in GPT-5.1?
At launch, OpenAI introduced optional prompt-cache retention of up to 24 hours for GPT-5.1, helping repeated requests reuse stable prompt prefixes for lower latency and cost.
Did GPT-5.1 support Apply Patch and shell workflows?
Yes. OpenAI introduced an Apply Patch tool and a shell workflow with GPT-5.1 for iterative code editing and command execution in agentic applications.
Can GPT-5.1 be fine-tuned?
No. The official model card lists fine-tuning as unsupported for GPT-5.1.
Model information
Last updated
The specifications, pricing, reasoning controls, lifecycle status, and historical feature details on this page are based on OpenAI's official GPT-5.1 model card, GPT-5.1 developer launch article, and API deprecation documentation. Provider behavior and availability may change, so migration decisions should be checked against the latest OpenAI documentation.