OpenAI GPT-5 model

GPT-5.1

A GPT-5 generation model built around efficient adaptive reasoning, responsive coding, agentic tool use, and a no-reasoning mode for latency-sensitive tasks.

Context window
400K
tokens
Max output
128K
tokens
Input
$1.25
per 1M tokens
Cached input
$0.125
per 1M tokens
Output
$10.00
per 1M tokens

01 / Overview

What GPT-5.1 Is

GPT-5.1 is an OpenAI model for coding and agentic tasks that introduced a more adaptive reasoning strategy: it can spend fewer tokens on straightforward work, think longer on harder problems, or disable reasoning entirely for latency-sensitive requests.

Efficient Reasoning Was the Core Design Goal

GPT-5.1 was released in November 2025 as the next step in the GPT-5 series, with a focus on balancing intelligence and speed across coding and agentic workloads.

Instead of spending a similar amount of reasoning effort on every request, GPT-5.1 was trained to adapt its thinking more dynamically to task complexity. Straightforward requests could complete with less hidden reasoning, while difficult tasks could remain persistent and explore more options before returning an answer.

The release also added a none reasoning mode. This made GPT-5.1 useful in architectures that wanted one model family to cover both fast tool-calling requests and deeper reasoning tasks.

OpenAI paired that reasoning design with stronger coding behavior, improved steerability, extended prompt caching, and new code-oriented tools such as Apply Patch and shell workflows.

GPT-5.1 is now deprecated. OpenAI announced deprecation on October 1, 2026, with API shutdown scheduled for April 1, 2027. GPT-6 Sol is the recommended replacement.

  • Designed for coding and agentic tasks.
  • Adaptive reasoning reduces unnecessary thinking on easier work.
  • Supports a no-reasoning mode for fast requests.
  • Introduced extended prompt caching up to 24 hours.
  • Deprecated with GPT-6 Sol as the recommended replacement.
Model profile
Provider
OpenAI
Family
GPT-5.1
Positioning
Coding & agentic
Status
Deprecated
Knowledge cutoff
Sep 30, 2024
Input modalities
Text, Image
Default reasoning
None

02 / Use cases

Where GPT-5.1 Fits Best

GPT-5.1 was strongest in systems that mixed fast routine actions with harder coding and agentic tasks, because one model could move from no-reasoning execution to deeper deliberation as workload complexity changed.

One Model Across Fast and Difficult Agent Steps

Many agent workflows contain a mixture of tasks.

One step may simply choose a tool, classify a request, or make a small code edit. Another may require debugging a failure across several files, planning a refactor, or reasoning through an ambiguous requirement.

GPT-5.1's reasoning controls were well suited to that pattern. Developers could use none or low for responsive operations, then allocate more inference effort to harder cases.

Coding was a major focus of the release. OpenAI described improvements in coding personality, steerability, code quality, front-end generation, and user-facing progress updates during sequences of tool calls.

The model also targeted tool-heavy agents. OpenAI specifically highlighted improved parallel tool calling in no-reasoning mode and introduced Apply Patch and shell workflows with GPT-5.1.

Because the model is deprecated, these use cases are now most useful as migration benchmarks for existing systems rather than reasons to start a new GPT-5.1 integration.

  • Responsive coding assistants with frequent short turns.
  • Agent workflows that mix simple and difficult steps.
  • Parallel tool calling at low reasoning effort.
  • Multi-file code editing and iterative repair.
  • Legacy GPT-5.1 workload benchmarking before migration.
Adaptive agent workloads
  1. 01

    Fast tool steps

    Use no-reasoning behavior for latency-sensitive actions, routing, and straightforward tool calls.

  2. 02

    Coding iteration

    Handle quick code edits, repository changes, frontend work, and interactive developer loops.

  3. 03

    Harder engineering

    Increase reasoning effort when debugging, planning, or multi-file changes require more persistence.

  4. 04

    Agent orchestration

    Combine function calls, shell execution, patches, search, and repeated model turns inside a controlled workflow.

03 / Pricing

GPT-5.1 Pricing

GPT-5.1 is priced at $1.25 per million input tokens, $0.125 per million cached input tokens, and $10.00 per million output tokens.

Reasoning Efficiency Was Part of the Cost Story

GPT-5.1's pricing matched GPT-5 at launch, but OpenAI positioned the newer model as more token-efficient because it could spend less time reasoning on easy tasks.

That distinction matters in agent systems. Two models with the same nominal token price can produce different end-to-end costs if one uses fewer reasoning tokens, finishes in fewer turns, or avoids unnecessary actions.

Cached input was priced at one tenth of standard input. GPT-5.1 also introduced extended prompt caching with retention of up to 24 hours, which made cached context more useful for long-running conversations, coding sessions, and retrieval workflows.

For migration analysis, preserve complete workflow measurements. Compare not just input and output prices, but also reasoning usage, number of turns, retries, tool calls, and completion rate.

  • $1.25 per 1M input tokens.
  • $0.125 per 1M cached input tokens.
  • $10.00 per 1M output tokens.
  • Extended prompt caching could retain eligible prefixes for up to 24 hours.
  • Evaluate total cost per completed agent or coding task.
Token pricing

1M tokens · USD

Input
$1.25
Cached input
$0.125
Output
$10.00

Example: 10K input + 2K output

Input cost
$0.0125
Output cost
$0.0200
Estimated total
$0.0325

04 / Context

A 400K Context Window for Agentic Work

GPT-5.1 provides a 400,000-token context window and supports up to 128,000 output tokens, giving coding and agent systems substantial room for instructions, source files, tool definitions, conversation state, and intermediate results.

Stable Prefixes Could Remain Cached Longer

The context window itself was large enough for significant coding and professional workloads, but GPT-5.1's more distinctive context feature was extended prompt caching.

At launch, OpenAI allowed developers to request cache retention of up to 24 hours. That was designed to improve latency and cost for follow-up requests that shared the same prompt prefix.

Coding sessions are a clear example. System instructions, repository guidance, tool definitions, and stable project context may remain unchanged while the user and agent iterate over several steps.

The same pattern appears in support agents, research workflows, and knowledge assistants.

A large context window still benefits from deliberate selection. Sending every available file or conversation turn can increase token usage and dilute the relevant signal even when everything technically fits.

  • 400,000-token context window.
  • Maximum output of 128,000 tokens.
  • Extended caching supported up to 24 hours at launch.
  • Useful for repeated coding and agent interactions with stable prefixes.
  • Context selection remains important for cost and relevance.
Context capacity

Context window

400,000

Max output

128,000

Input contextOutput limit

A GPT-5.1 coding session could reuse stable system instructions, tool schemas, repository guidance, and project context while new user requests and tool results changed between turns.

05 / Reasoning

Adaptive Reasoning from None to High

GPT-5.1 supports none, low, medium, and high reasoning effort, with none as the default.

No Reasoning Was a First-Class Operating Mode

The none setting was one of the defining additions in GPT-5.1.

OpenAI designed it for latency-sensitive tasks that did not require deep deliberation. In this mode, the model behaved more like a non-reasoning model while retaining GPT-5.1's coding, instruction-following, and tool-calling capabilities.

For more complex requests, developers could move to low, medium, or high.

OpenAI's launch guidance recommended low or medium for higher-complexity tasks and high when intelligence and reliability mattered more than speed.

The model also adapted its thinking within a chosen reasoning level. OpenAI described GPT-5.1 as spending fewer reasoning tokens on straightforward tasks while remaining persistent on harder tasks.

Reasoning effort
none · defaultlowmediumhigh

Latency-sensitive step

Use none for simple tool calls, quick edits, routing, and other tasks where extra deliberation does not improve the result.

Hard agent task

Increase to low, medium, or high when debugging, planning, research, or multi-step execution requires more reliability.

06 / Capabilities

Coding, Tool Calling, Patching, and Shell Workflows

GPT-5.1 supports streaming, function calling, structured outputs, text and image input, and was launched with new Apply Patch and shell workflows designed for iterative agentic coding.

A Model Designed to Do More Than Return Code

The official model card lists streaming, function calling, and structured outputs as supported.

GPT-5.1 also accepts image input, which can be useful when coding or agent tasks include screenshots, visual references, diagrams, or interface state.

At launch, OpenAI introduced Apply Patch with GPT-5.1. The tool allowed the model to create, update, and delete files using structured diffs that an application could apply and report back on.

OpenAI also introduced a shell workflow that allowed GPT-5.1 to propose commands, receive execution results, and continue through a plan-execute loop.

Web search was supported with no-reasoning mode as well, and the release emphasized stronger parallel tool calling.

Because GPT-5.1 is deprecated, current tool availability should be checked before relying on these historical integrations, and new systems should target GPT-6 Sol instead.

  • Streaming supported.
  • Function calling supported.
  • Structured outputs supported.
  • Text and image input with text output.
  • Apply Patch introduced with GPT-5.1.
  • Shell workflows introduced with GPT-5.1.
  • Web search supported in documented GPT-5.1 workflows.
  • Fine-tuning and predicted outputs are not supported.
Coding and agent capabilities
  • Function calling

    Connect model decisions to application-defined actions and parallel tool workflows.

    Supported
  • Structured outputs

    Return machine-readable responses that conform to an application schema.

    Supported
  • Apply Patch

    Historically introduced with GPT-5.1 for structured create, update, and delete operations across code files.

    Supported
  • Shell workflow

    Historically introduced with GPT-5.1 for command execution loops where the integration runs commands and returns results.

    Supported
  • Image input

    Use screenshots and other visual context alongside text and code.

    Supported
  • Fine-tuning

    OpenAI lists fine-tuning as unsupported for GPT-5.1.

    Not listed

07 / Evaluation

Strengths and Limitations

GPT-5.1 represented an important shift toward more efficient adaptive reasoning for coding and agents, but its deprecated lifecycle now makes it primarily a migration and regression-testing target rather than a model for new production systems.

Historical strengths

  • Adaptive reasoning

    GPT-5.1 was trained to spend less reasoning effort on easier tasks while remaining persistent when the task became difficult.

  • True no-reasoning mode

    The none setting gave developers a low-latency operating mode inside the same model family used for deeper agentic work.

  • Coding-focused improvements

    OpenAI emphasized better steerability, code quality, frontend generation, tool-call updates, and iterative developer experience.

  • Extended prompt caching

    Up to 24-hour cache retention improved the economics of repeated multi-turn sessions with stable prompt prefixes.

What to consider

  • Deprecated lifecycle

    GPT-5.1 was deprecated on October 1, 2026 and is scheduled to shut down on April 1, 2027. OpenAI recommends GPT-6 Sol.

  • Older knowledge cutoff

    The official model card lists September 30, 2024 as the knowledge cutoff.

  • No xhigh reasoning

    GPT-5.1 supports none, low, medium, and high, while later GPT generations extend the reasoning range further.

  • Newer agent models supersede it

    Current OpenAI models provide newer capabilities, longer support horizons, and updated tool ecosystems for coding and agent workloads.

Preserve behavior before migration

Benchmark GPT-5.1 before moving your coding and agent workloads

Capture representative coding, tool-calling, low-latency, and long-context cases, then compare quality, reasoning behavior, token usage, and cost against GPT-6 Sol in EidoStack.

Start Free

GPT-5.1 is deprecated. OpenAI recommends migrating production workloads to GPT-6 Sol before the April 1, 2027 API shutdown.

Common Questions

What is GPT-5.1?

GPT-5.1 is an OpenAI model designed for coding and agentic tasks with configurable reasoning. It introduced more adaptive reasoning behavior, a no-reasoning mode, extended prompt caching, and new coding-oriented tool workflows.

Is GPT-5.1 deprecated?

Yes. OpenAI deprecated GPT-5.1 on October 1, 2026 and plans to remove it from the API on April 1, 2027.

What should replace GPT-5.1?

OpenAI lists GPT-6 Sol as the recommended replacement for GPT-5.1.

How much does GPT-5.1 cost?

OpenAI lists $1.25 per 1M input tokens, $0.125 per 1M cached input tokens, and $10.00 per 1M output tokens.

What is the context window of GPT-5.1?

GPT-5.1 has a 400,000-token context window and supports up to 128,000 output tokens.

What reasoning levels does GPT-5.1 support?

GPT-5.1 supports none, low, medium, and high reasoning effort. None is the default.

What is adaptive reasoning in GPT-5.1?

OpenAI trained GPT-5.1 to vary how much effort it spends thinking based on task complexity. Straightforward tasks can use fewer reasoning tokens, while harder tasks can remain more persistent.

What is GPT-5.1 no-reasoning mode?

Setting reasoning effort to none disables extended reasoning for latency-sensitive tasks while retaining GPT-5.1's general intelligence, coding, and tool-calling behavior.

What is the knowledge cutoff for GPT-5.1?

OpenAI lists September 30, 2024 as the knowledge cutoff for GPT-5.1.

Does GPT-5.1 support image input?

Yes. GPT-5.1 accepts text and image input and produces text output. Direct audio and video modalities are not supported.

Does GPT-5.1 support structured outputs?

Yes. The official model card lists structured outputs, function calling, and streaming as supported.

What was extended prompt caching in GPT-5.1?

At launch, OpenAI introduced optional prompt-cache retention of up to 24 hours for GPT-5.1, helping repeated requests reuse stable prompt prefixes for lower latency and cost.

Did GPT-5.1 support Apply Patch and shell workflows?

Yes. OpenAI introduced an Apply Patch tool and a shell workflow with GPT-5.1 for iterative code editing and command execution in agentic applications.

Can GPT-5.1 be fine-tuned?

No. The official model card lists fine-tuning as unsupported for GPT-5.1.

Model information

Last updated

The specifications, pricing, reasoning controls, lifecycle status, and historical feature details on this page are based on OpenAI's official GPT-5.1 model card, GPT-5.1 developer launch article, and API deprecation documentation. Provider behavior and availability may change, so migration decisions should be checked against the latest OpenAI documentation.

GPT-5.1 — Pricing, 400K Context & Adaptive Reasoning | EidoStack