OpenAI frontier model

GPT-5.4

A general-purpose frontier model for professional work across coding, document analysis, tool-heavy agents, computer use, and long-context workflows.

Context window
1.05M
tokens
Max output
128K
tokens
Input
$2.50
per 1M tokens
Cached input
$0.25
per 1M tokens
Output
$15.00
per 1M tokens

01 / Overview

What GPT-5.4 Is

GPT-5.4 is an OpenAI frontier model for complex professional work, designed to move between software engineering, reasoning, writing, document analysis, multimodal input, and tool use without requiring a separate specialist model for each stage.

A General-Purpose Frontier Model for Real Work

OpenAI introduced GPT-5.4 as a model for professional work across the API and Codex. Its role is deliberately broad: developers can use it to analyze complex information, build production software, work with large collections of documents, and automate workflows that require several model and tool interactions.

Compared with earlier GPT-5 models, GPT-5.4 improved coding, instruction following, document understanding, image perception, tool use, and long-running task execution. OpenAI also emphasizes better token efficiency across tool-heavy workflows, where fewer unnecessary calls can matter as much as the price of each individual token.

GPT-5.4 was also an important architectural step in the model family. It introduced built-in computer use to the mainline model series, native compaction for longer agent trajectories, and Tool Search for large tool ecosystems.

  • General-purpose frontier model for complex professional work.
  • Strong software-engineering and code-heavy workflow support.
  • Built for document-heavy and multi-step agentic tasks.
  • First mainline GPT model with built-in computer-use support.
  • Supports Tool Search and native compaction for long-running workflows.
Model profile
Provider
OpenAI
Family
GPT-5.4
Positioning
Professional frontier model
Knowledge cutoff
Aug 31, 2025
Input modalities
Text, Image
Output modality
Text
Default reasoning
None

02 / Use cases

Where GPT-5.4 Fits Best

GPT-5.4 is most useful when a workload crosses boundaries: code and documents, reasoning and tools, structured data and written output, or analysis and direct interaction with software.

One Model Across Several Stages of Work

A software task may begin with a product requirement, continue through repository analysis and implementation, and finish with verification. GPT-5.4 is designed to participate across that sequence instead of acting only as a code generator.

The same flexibility matters in business workflows. OpenAI highlights document-heavy and spreadsheet-heavy use cases in analytics, finance, and customer service. A model can inspect source material, reason over the data, use tools to obtain additional context, and generate a structured or polished result.

Multimodal input broadens the workload further. Images can be included alongside text, making GPT-5.4 relevant for workflows that need to interpret screenshots, diagrams, scanned material, or visual application state.

For agent systems, the key value is continuity across steps. Tool Search, computer use, shell access, file retrieval, and code-oriented tools allow the model to move from understanding a task toward actually completing it.

  • Production software development and multi-file coding work.
  • Document-heavy analysis and business workflows.
  • Spreadsheet-oriented analytics and finance tasks.
  • Tool-driven agents that search, execute, inspect, and revise.
  • Multimodal workflows that combine text with visual input.
Professional workflows
  1. 01

    Software engineering

    Build, modify, debug, and reason about production software while following repository-specific patterns.

  2. 02

    Document and data analysis

    Work across long documents, spreadsheets, reports, and business source material.

  3. 03

    Computer-use automation

    Interact with software interfaces as part of a build, run, verify, and fix workflow.

  4. 04

    Tool-heavy agents

    Search for relevant tools, load only needed definitions, execute actions, and continue across multiple steps.

03 / Pricing

GPT-5.4 Pricing

GPT-5.4 standard pricing is $2.50 per million input tokens, $0.25 per million cached input tokens, and $15.00 per million output tokens.

Evaluate End-to-End Workflow Cost

For GPT-5.4, cost depends heavily on workflow shape. A simple single-turn request may be dominated by output tokens, while an agentic or document-heavy task can accumulate substantial input across source files, tool results, and conversation history.

Prompt caching can reduce the price of repeated stable prefixes. GPT-5.4 belongs to OpenAI's pre-GPT-5.6 caching generation: caching is automatic for eligible prefixes, a stable prompt_cache_key can help optimize routing, and there is no additional cache-write charge.

Agentic efficiency matters as well. OpenAI reports improvements in end-to-end performance and token efficiency on tool-heavy workflows compared with earlier models. Fewer unnecessary tool calls and shorter trajectories can reduce total workflow cost even when headline token rates stay fixed.

Very long prompts need separate budgeting. Once input exceeds 272K tokens, OpenAI applies higher long-context rates to the full session.

  • $2.50 per 1M standard input tokens.
  • $0.25 per 1M cached input tokens.
  • $15.00 per 1M output tokens.
  • No additional prompt-cache write charge.
  • Tool-specific operations can introduce separate fees.
Token pricing

1M tokens · USD

Input
$2.50
Cached input
$0.25
Output
$15.00

Example: 10K input + 2K output

Input cost
$0.0250
Output cost
$0.0300
Estimated total
$0.0550

04 / Context

A 1.05M-Token Context for Codebases and Document Collections

GPT-5.4 provides a 1,050,000-token context window and supports up to 128,000 output tokens, making it suitable for workloads that need to reason across large codebases, long document collections, or extended agent trajectories.

Large Context Changes How Workflows Can Be Designed

A large context window allows related information to remain together instead of being aggressively split into isolated model calls.

For software engineering, that can include source files, repository conventions, issue descriptions, logs, and implementation history. For professional analysis, it can include multiple reports, spreadsheet data, retrieved files, and instructions that define how the final result should be structured.

GPT-5.4 also introduced native compaction support. Compaction is useful for long-running agent workflows because it helps preserve important context while reducing the amount of accumulated history that needs to remain verbatim in later steps.

This does not eliminate the need for context discipline. Sending irrelevant material still increases cost and can make important signals harder to identify. Above 272K input tokens, long-context pricing also becomes materially higher.

  • 1,050,000-token context window.
  • Maximum output of 128,000 tokens.
  • Suitable for large codebases and long document collections.
  • Native compaction supports longer agent trajectories.
  • Requests above 272K input tokens use higher pricing.
Context capacity

Context window

1,050,000

Max output

128,000

Input contextOutput limit

Context can include system instructions, code, documents, spreadsheets, retrieved files, conversation history, tool results, and other working state required by the task.

05 / Reasoning

Reasoning from None to XHigh

GPT-5.4 supports none, low, medium, high, and xhigh reasoning effort. Unlike several later reasoning-oriented models, none is the default.

Start with the Lowest Effort That Meets the Quality Bar

The default none setting makes GPT-5.4 flexible for applications that want strong general-purpose behavior without automatically paying for deeper reasoning on every request.

More difficult workloads can move upward through low, medium, high, or xhigh. OpenAI's migration guidance recommends medium or high reasoning for workloads that previously depended on dedicated reasoning models such as o3, while tasks coming from GPT-4.1 can often begin at none.

This creates a practical evaluation strategy: keep a representative test set, establish the quality baseline at none, then increase reasoning only where the workload demonstrates a measurable benefit.

For professional applications, this is particularly useful because coding, document processing, and business analysis do not all need the same inference depth.

Reasoning effort
none · defaultlowmediumhighxhigh

General professional task

Start at none for well-specified writing, coding, extraction, and structured application work.

Complex reasoning task

Increase to medium, high, or xhigh when analysis, planning, or multi-step decisions require additional reasoning depth.

06 / Capabilities

Capabilities for Modern Agent Systems

GPT-5.4 supports text and image input, streaming, function calling, structured outputs, computer use, Tool Search, and a broad set of Responses API tools for building production agents.

Tool Search and Computer Use Expanded the Mainline Model

GPT-5.4 introduced Tool Search to OpenAI's mainline model family. Tool Search allows large tool ecosystems to defer definitions and load only the tools that are relevant to the current task, reducing prompt size and improving tool selection.

The model was also the first mainline GPT model with built-in computer-use support. That enables workflows in which an agent interacts directly with software, observes the result, and continues through a build-run-verify-fix loop.

Other supported tools include web search, file search, image generation, code interpreter, hosted shell, Apply Patch, Skills, MCP, and Tool Search itself.

Function calling and structured outputs provide the application-level control needed to connect these capabilities to production systems. Fine-tuning is currently not supported.

  • Text input and output with image input.
  • Streaming responses.
  • Function calling and structured outputs.
  • Web search and file search.
  • Image generation and code interpreter.
  • Hosted shell and Apply Patch.
  • Skills and built-in computer use.
  • MCP and Tool Search.
  • Fine-tuning is currently not supported.
Supported capabilities
  • Tool Search

    Search large tool ecosystems and load only relevant tool definitions when needed.

    Supported
  • Computer use

    Interact directly with software as part of multi-step automated workflows.

    Supported
  • Coding tools

    Use code interpreter, hosted shell, and Apply Patch for technical execution.

    Supported
  • Search and retrieval

    Use web search and file search to ground work in external information.

    Supported
  • Structured integration

    Use function calling and structured outputs for application-controlled results.

    Supported
  • Multimodal input

    Analyze text and images within the same workflow.

    Supported

07 / Evaluation

Strengths and Limitations

GPT-5.4 is a versatile professional model with unusually broad workflow coverage, but later model generations may offer better economics or capability for specialized production roles.

Strengths

  • Broad professional coverage

    One model can move between coding, reasoning, writing, documents, multimodal input, and tool-driven workflows.

  • Strong agent foundations

    Computer use, Tool Search, native compaction, shell access, patching, and search support substantial multi-step automation.

  • Large context window

    A 1.05M-token context supports codebases, long document collections, business data, and extended agent trajectories.

  • Flexible reasoning control

    Reasoning ranges from none through xhigh, allowing applications to vary inference effort across task classes.

What to consider

  • Later generations are available

    GPT-5.5, GPT-5.6, and GPT-6-family models provide newer options that may be preferable for some workloads.

  • Long-context pricing multiplier

    Prompts above 272K input tokens are billed at 2× input and 1.5× output pricing across the full session.

  • Earlier caching architecture

    GPT-5.4 uses pre-GPT-5.6 prompt caching with automatic breakpoints and prompt_cache_key routing rather than explicit cache breakpoints.

  • Fine-tuning is unavailable

    The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.4.

Evaluate the full workflow

Test GPT-5.4 on your real professional tasks

Run representative coding, document, multimodal, and tool-driven workloads with your actual prompts and context, then inspect quality, token usage, cost, and context consumption in EidoStack.

Start Free

Connect your own OpenAI API key and evaluate GPT-5.4 under the same reasoning, context, and tool conditions your application will use.

Common Questions

What is GPT-5.4?

GPT-5.4 is an OpenAI frontier model for complex professional work. It is designed for general-purpose workflows spanning coding, reasoning, document analysis, writing, multimodal input, and tool use.

How much does GPT-5.4 cost?

Standard pricing is $2.50 per 1M input tokens, $0.25 per 1M cached input tokens, and $15.00 per 1M output tokens. Tool-specific charges may apply separately.

What is the context window of GPT-5.4?

GPT-5.4 has a 1,050,000-token context window and supports up to 128,000 output tokens.

What reasoning levels does GPT-5.4 support?

GPT-5.4 supports none, low, medium, high, and xhigh reasoning effort. None is the default.

What is the knowledge cutoff for GPT-5.4?

OpenAI lists August 31, 2025 as the knowledge cutoff for GPT-5.4.

Does GPT-5.4 support image input?

Yes. GPT-5.4 accepts text and image input and generates text output. Direct audio and video modalities are not supported.

What tools does GPT-5.4 support?

OpenAI currently lists web search, file search, image generation, code interpreter, hosted shell, Apply Patch, Skills, computer use, MCP, and Tool Search as supported Responses API tools.

What is Tool Search in GPT-5.4?

Tool Search allows a large tool ecosystem to defer tool definitions and load only the relevant tools when the model needs them. This can reduce token usage and improve tool selection in production agents.

Does GPT-5.4 support computer use?

Yes. GPT-5.4 was the first mainline OpenAI model with built-in computer-use support, allowing agents to interact with software during multi-step workflows.

Does GPT-5.4 support prompt caching?

Yes. GPT-5.4 uses OpenAI's earlier automatic prompt-caching architecture. Cached input is billed at $0.25 per 1M tokens and there is no additional cache-write charge.

What workloads are a good fit for GPT-5.4?

GPT-5.4 is well suited to production software engineering, document and spreadsheet analysis, multimodal business workflows, tool-heavy agents, computer-use automation, and long-context professional tasks.

Can GPT-5.4 be fine-tuned?

No. The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.4.

Model information

Last updated

The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.4. Provider pricing, model limits, prompt-caching behavior, processing tiers, tools, and availability may change, so production assumptions should be checked against the latest provider documentation.

GPT-5.4 — Pricing, Context Window & Agentic Capabilities | EidoStack