OpenAI GPT-5 model

GPT-5

OpenAI's original GPT-5 API flagship for coding, reasoning, and agentic tasks, introducing minimal reasoning, response verbosity controls, and custom tools for developer workflows.

Context window
400K
tokens
Max output
128K
tokens
Input
$1.25
per 1M tokens
Cached input
$0.125
per 1M tokens
Output
$10.00
per 1M tokens

01 / Overview

What GPT-5 Is

GPT-5 is OpenAI's original GPT-5 API flagship for coding, reasoning, and agentic tasks, notable for combining configurable reasoning with strong tool use, long-context retrieval, response-length control, and developer-oriented agent behavior.

The First GPT-5 Developer Flagship

OpenAI released GPT-5 in the API in August 2025 as its strongest model at the time for coding and agentic tasks.

The model was designed to act less like a one-shot text generator and more like a coding collaborator. OpenAI emphasized bug fixing, code editing, complex-codebase questions, front-end generation, detailed instruction following, and sustained work across multiple tool calls.

GPT-5 also introduced several controls that shaped later OpenAI models. reasoning_effort gained a minimal setting for faster answers. A new verbosity parameter let developers steer the default response length. Custom tools allowed free-form plaintext tool inputs instead of requiring every call to use JSON.

Agentic behavior was another defining theme. OpenAI highlighted the model's ability to chain many tool calls, recover from tool errors, use parallel calls, and provide visible progress messages before and between actions.

The model is now deprecated. OpenAI's current model card recommends GPT-6 Astra for new flagship workloads.

  • Original GPT-5 API flagship for coding and agents.
  • Introduced minimal reasoning effort.
  • Introduced the verbosity response-control parameter.
  • Introduced custom tools with free-form tool input.
  • Designed for sustained multi-step tool execution.
Model profile
Provider
OpenAI
Family
GPT-5
Positioning
Original GPT-5 flagship
Status
Deprecated
Knowledge cutoff
Sep 30, 2024
Input modalities
Text, Image
Output modality
Text

02 / Use cases

Where GPT-5 Fit Best

GPT-5 was built for coding and agent workflows where the model needed to understand a substantial task, follow detailed instructions, call several tools, and continue working until it reached a useful result.

From Code Generation to Coding Collaboration

Coding was one of GPT-5's primary launch use cases.

OpenAI trained the model to fix bugs, edit existing code, answer questions about complex repositories, and produce front-end implementations. The launch material emphasized not only code quality but also collaboration: GPT-5 could provide plans, progress updates, and recaps while working through tool calls.

Long-running agents were another major target. Instead of completing one tool invocation and stopping, GPT-5 was designed to chain actions in sequence or parallel, inspect results, handle failures, and continue toward the original goal.

The model was also useful for tasks where reasoning depth varied. A straightforward retrieval or transformation could run with minimal reasoning, while difficult coding or decision-heavy work could use a higher effort level.

Today these workloads are most useful as migration benchmarks. Existing teams can capture how GPT-5 behaved on representative engineering and agent tasks, then compare newer models against the same evaluation set.

  • Repository-level bug fixing and code editing.
  • Front-end generation and iterative implementation.
  • Long-running agents with several tool calls.
  • Complex instruction-following workflows.
  • Migration regression testing for legacy GPT-5 applications.
Original GPT-5 workloads
  1. 01

    Coding collaboration

    Plan changes, inspect code, edit files, and communicate progress during longer engineering tasks.

  2. 02

    Tool orchestration

    Chain application tools sequentially or in parallel and continue after observing intermediate results.

  3. 03

    Complex codebase work

    Answer questions, locate relevant logic, fix bugs, and reason about changes across a substantial repository context.

  4. 04

    Agentic execution

    Follow detailed instructions across multiple steps instead of treating every request as a single response.

03 / Pricing

GPT-5 Pricing

GPT-5 is listed at $1.25 per million input tokens, $0.125 per million cached input tokens, and $10.00 per million output tokens.

Measure Agent Cost Across the Whole Task

GPT-5's token prices are only one part of the economics for an agentic workflow.

A coding task may involve a large initial prompt, several reasoning phases, tool calls, intermediate outputs, and follow-up turns. A model that completes the task in fewer actions can be cheaper overall even if its per-token price is similar to another model.

Cached input is one tenth of standard input pricing, which can help when repeated requests share a stable prefix such as system instructions, tool definitions, repository rules, or project context.

Reasoning tokens are part of generated usage and therefore affect output-side cost. Higher reasoning effort can improve difficult tasks while increasing latency and token consumption.

Because GPT-5 is deprecated, these rates are now most useful as a historical baseline for migration analysis.

  • $1.25 per 1M input tokens.
  • $0.125 per 1M cached input tokens.
  • $10.00 per 1M output tokens.
  • Reasoning usage contributes to generated-token cost.
  • Tool-specific services may add separate fees.
Token pricing

1M tokens · USD

Input
$1.25
Cached input
$0.125
Output
$10.00

Example: 10K input + 2K output

Input cost
$0.0125
Output cost
$0.0200
Estimated total
$0.0325

04 / Context

A 400K Context Window for Code and Agent State

GPT-5 provides a 400,000-token context window and supports up to 128,000 output tokens, giving the model room for substantial source code, instructions, retrieved content, conversation state, and tool results.

Long Context Supported More Ambitious Tasks

Large coding tasks rarely depend on one file.

A useful working set may include repository instructions, source code, interfaces, tests, issue descriptions, documentation, earlier decisions, and command output. GPT-5's 400K window made it possible to keep much of that material available during one workflow.

OpenAI also highlighted long-context retrieval as an area of strength at launch. This mattered for agent systems that needed to locate relevant details inside large inputs rather than simply summarize them.

Still, a large window does not eliminate the need for context discipline. Irrelevant files and repeated history add cost and can make the important signal harder to identify.

Migration testing should therefore preserve context strategy as a separate variable. First compare models on the same working set; then optimize retrieval or history management independently.

  • 400,000-token context window.
  • Maximum output of 128,000 tokens.
  • Suitable for large code and agent working sets.
  • Designed for strong long-context retrieval.
  • Context selection remains important for cost and accuracy.
Context capacity

Context window

400,000

Max output

128,000

Input contextOutput limit

A GPT-5 agent context can contain repository instructions, selected source files, tests, user requirements, tool definitions, retrieved information, and intermediate execution results.

05 / Reasoning

Minimal to High Reasoning

GPT-5 supports minimal, low, medium, and high reasoning effort, allowing developers to trade response speed against deeper deliberation.

Minimal Reasoning Was a New Developer Control

GPT-5 introduced minimal as a reasoning-effort option.

The goal was to reduce thinking time for requests that did not benefit from extensive deliberation. OpenAI positioned higher effort levels as the quality-oriented end of the range and lower values as the speed-oriented end.

At launch, medium was the default reasoning level, while minimal, low, and high gave developers explicit alternatives.

This mattered because different agent steps have different requirements. A simple retrieval or tool selection may not benefit from high reasoning, while a difficult debugging problem can justify more inference work.

GPT-5 therefore helped establish a pattern that became central to later OpenAI models: evaluate reasoning effort per workload rather than treating maximum reasoning as the universal default.

Reasoning effort
minimallowmedium · launch defaulthigh

Fast agent step

Use minimal or low for straightforward retrieval, simple tool selection, and low-ambiguity transformations.

Difficult reasoning task

Use medium or high when coding, debugging, planning, or multi-step execution benefits from additional deliberation.

06 / Capabilities

Verbosity, Custom Tools, and Structured Agent Output

GPT-5 supports streaming, function calling, structured outputs, text and image input, and introduced developer controls such as response verbosity and free-form custom tools.

More Control Over How the Model Responds and Acts

The official model card lists streaming, function calling, and structured outputs as supported.

GPT-5 also accepts image input, allowing visual information to participate in coding, reasoning, and agent workflows.

A distinctive launch feature was verbosity. Developers could choose low, medium, or high to influence the default amount of detail in the final response while still allowing explicit prompt instructions to take precedence.

Custom tools were another important addition. Instead of requiring developer-defined tool arguments to always be valid JSON, GPT-5 could call custom tools using plaintext. OpenAI also allowed developers to constrain these free-form tool calls with grammars.

The launch release additionally highlighted parallel tool calling and built-in tools such as web search, file search, and image generation.

Fine-tuning and predicted outputs are not supported on the current GPT-5 model card.

  • Streaming supported.
  • Function calling supported.
  • Structured outputs supported.
  • Text and image input with text output.
  • verbosity supports low, medium, and high.
  • Custom tools support free-form text input.
  • Parallel and built-in tool workflows were part of the GPT-5 launch.
  • Fine-tuning and predicted outputs are not supported.
Developer controls
  • Verbosity control

    Steer the default response length with low, medium, or high verbosity.

    Supported
  • Custom tools

    Call developer-defined tools with free-form plaintext instead of requiring JSON-only arguments.

    Supported
  • Structured outputs

    Return machine-readable responses that conform to an application schema.

    Supported
  • Parallel tool calling

    Coordinate multiple tool operations as part of longer agentic workflows.

    Supported
  • Image input

    Include screenshots, diagrams, and other visual context alongside text.

    Supported
  • Fine-tuning

    OpenAI currently lists fine-tuning as unsupported for GPT-5.

    Not listed

07 / Evaluation

Strengths and Limitations

GPT-5 was a major developer model for coding and tool-driven agents, but its current value is primarily historical and migratory: newer model families provide a longer support horizon, newer reasoning controls, and more capable tool ecosystems.

Historical strengths

  • Strong agentic coding focus

    OpenAI designed GPT-5 around real coding collaboration, repository work, bug fixing, frontend generation, and sustained engineering tasks.

  • Reliable tool orchestration

    The model was trained to chain many tool calls, use parallel actions, recover from tool errors, and maintain progress toward a larger goal.

  • Flexible reasoning depth

    Minimal through high reasoning allowed developers to tune speed and deliberation for different task classes.

  • Developer response controls

    Verbosity and custom tools provided direct control over response length and tool-call formats.

What to consider

  • Deprecated model

    OpenAI's current model card marks GPT-5 as deprecated and recommends GPT-6 Astra for new flagship workloads.

  • Deprecated dated snapshot

    The gpt-5-2025-08-07 snapshot is scheduled for shutdown on December 11, 2026, with GPT-5.6 Sol listed as its replacement.

  • Older knowledge cutoff

    The model card lists September 30, 2024 as the knowledge cutoff.

  • Newer models offer broader controls

    Later GPT generations add newer reasoning modes, larger context windows, expanded tool support, and more current model knowledge.

Preserve your GPT-5 baseline

Benchmark GPT-5 behavior before migrating

Run representative coding, tool-calling, reasoning, and long-context workloads, then compare quality, latency, token usage, and cost against current OpenAI models in EidoStack.

Start Free

GPT-5 is deprecated. Use historical evaluations to preserve expected behavior while moving production workloads to a supported model.

Common Questions

What is GPT-5?

GPT-5 is OpenAI's original GPT-5 API flagship for coding, reasoning, and agentic tasks. It introduced minimal reasoning, response verbosity controls, custom tools, and stronger long-running tool orchestration.

Is GPT-5 deprecated?

Yes. OpenAI's current GPT-5 model card marks the model as deprecated and recommends GPT-6 Astra for new flagship API workloads.

When is the GPT-5 snapshot being shut down?

OpenAI's deprecation schedule lists December 11, 2026 as the shutdown date for the gpt-5-2025-08-07 snapshot.

What replaces the deprecated GPT-5 snapshot?

OpenAI's deprecation schedule lists GPT-5.6 Sol as the recommended replacement for gpt-5-2025-08-07, while the current GPT-5 model card recommends GPT-6 Astra as the latest flagship.

How much does GPT-5 cost?

The official model card lists $1.25 per 1M input tokens, $0.125 per 1M cached input tokens, and $10.00 per 1M output tokens.

What is the context window of GPT-5?

GPT-5 has a 400,000-token context window and supports up to 128,000 output tokens.

What reasoning levels does GPT-5 support?

GPT-5 supports minimal, low, medium, and high reasoning effort. OpenAI's original launch guidance described medium as the default.

What is minimal reasoning in GPT-5?

Minimal reasoning reduces the amount of deliberation before the answer, making it useful for tasks where lower latency matters more than deeper inference.

What is the GPT-5 verbosity parameter?

GPT-5 introduced a verbosity parameter with low, medium, and high settings to control the default level of detail in responses.

What are GPT-5 custom tools?

Custom tools allow GPT-5 to send free-form plaintext to a developer-defined tool instead of requiring all tool arguments to be encoded as JSON. Developers can also constrain these tool inputs with grammars.

Does GPT-5 support image input?

Yes. GPT-5 accepts text and image input and generates text output. Direct audio and video modalities are not supported.

Does GPT-5 support structured outputs?

Yes. The current model card lists structured outputs, function calling, and streaming as supported.

Can GPT-5 be fine-tuned?

No. The current OpenAI model card lists fine-tuning as unsupported for GPT-5.

Why keep a GPT-5 page after deprecation?

The page remains useful for migration research, historical pricing and capability comparisons, regression testing, and understanding applications originally built around the first GPT-5 API generation.

Model information

Last updated

The specifications, pricing, capabilities, and lifecycle information on this page are based on OpenAI's official GPT-5 model card, GPT-5 developer launch article, API changelog, and deprecation documentation. The current model card marks GPT-5 as deprecated and recommends GPT-6 Astra; OpenAI's deprecation schedule lists December 11, 2026 as the shutdown date for the gpt-5-2025-08-07 snapshot.

GPT-5 — Pricing, 400K Context & Agentic Coding | EidoStack