OpenAI mini model
GPT-5.4 Mini
OpenAI's strongest mini model for coding, computer use, and subagents, designed for fast, efficient execution across high-volume technical workflows.
- Context window
- 400K
- tokens
- Max output
- 128K
- tokens
- Input
- $0.75
- per 1M tokens
- Cached input
- $0.075
- per 1M tokens
- Output
- $4.50
- per 1M tokens
01 / Overview
What GPT-5.4 Mini Is
GPT-5.4 Mini is OpenAI's strongest mini model for coding, computer use, and subagents, built to deliver GPT-5.4-class capabilities in a faster and more efficient package for high-volume technical workloads.
Smaller Model, Execution-Focused Role
GPT-5.4 Mini is not positioned as a generic lightweight text model. OpenAI designed it around workflows where speed directly affects product quality: responsive coding assistants, parallel subagents, interface automation, and multimodal applications that need to interpret visual state quickly.
The model sits between a flagship model and a very small routing or extraction model. It can reason, use tools, inspect images, operate computer interfaces, and handle substantial coding work, while costing materially less than full GPT-5.4.
That makes it especially useful inside multi-model systems. A larger model can plan or coordinate a difficult task, then delegate narrower pieces of work to several GPT-5.4 Mini subagents running in parallel.
OpenAI also describes Mini as more literal than larger models. It performs well when the task is clearly structured, but developers should specify execution order, side effects, and expected outputs explicitly instead of assuming the model will infer missing workflow details.
- Strongest GPT-5.4 mini-tier model for coding, computer use, and subagents.
- Designed for fast, high-volume technical execution.
- Supports text and image input with text output.
- Includes reasoning from
nonethroughxhigh. - Supports a broad Responses API toolset.
- Provider
- OpenAI
- Family
- GPT-5.4
- Tier
- Mini
- Knowledge cutoff
- Aug 31, 2025
- Input modalities
- Text, Image
- Output modality
- Text
- Default reasoning
- None
02 / Use cases
Where GPT-5.4 Mini Fits Best
GPT-5.4 Mini is strongest when the application needs a capable model to execute many technical subtasks quickly: edit code, navigate repositories, inspect screenshots, call tools, or work as one of several parallel subagents.
Delegate Work Without Paying Flagship Cost for Every Step
Subagent architectures are one of the clearest use cases for GPT-5.4 Mini. A larger model can own planning and final evaluation while Mini agents perform bounded supporting tasks such as searching a codebase, reviewing large files, checking documentation, or preparing candidate changes.
Coding assistants are another strong fit. OpenAI highlights targeted edits, codebase navigation, frontend generation, and debugging loops where fast iteration matters. These are tasks where model latency can directly affect how usable the development experience feels.
Computer-use workflows also benefit from Mini's speed. The model can interpret screenshots and interact with interfaces, making it relevant for systems that repeatedly observe UI state and take actions.
The common pattern is frequent execution with enough complexity to require real reasoning and tools, but not so much ambiguity that every step needs flagship-level inference.
- Parallel coding subagents.
- Repository search and codebase navigation.
- Fast edit-debug-verify development loops.
- Computer-use and screenshot-driven automation.
- High-volume tool-calling workflows.
- 01
Coding assistants
Handle targeted edits, repository navigation, frontend work, and debugging loops with responsive execution.
- 02
Parallel subagents
Delegate narrow tasks such as code search, file review, documentation analysis, and supporting research.
- 03
Computer use
Interpret screenshots and interact with software interfaces when fast visual feedback matters.
- 04
Tool-heavy automation
Call search, file, shell, code, MCP, and application tools across repeated production workflows.
03 / Pricing
GPT-5.4 Mini Pricing
GPT-5.4 Mini standard pricing is $0.75 per million input tokens, $0.075 per million cached input tokens, and $4.50 per million output tokens.
Economics Designed for Repeated Technical Work
The main pricing advantage of Mini appears when an application runs many model calls rather than a few isolated prompts.
A subagent system can generate several parallel requests for a single user task. A coding assistant may repeatedly inspect files, propose changes, review tool results, and iterate. A computer-use workflow may observe and act many times before completion.
At those volumes, the gap between Mini and the full GPT-5.4 model can materially affect cost per completed workflow.
Cached input is priced at one tenth of the standard input rate, which can help when repeated requests share stable system instructions, repository guidance, tool definitions, or other reusable prefixes.
- $0.75 per 1M input tokens.
- $0.075 per 1M cached input tokens.
- $4.50 per 1M output tokens.
- Tool-specific operations may carry separate fees.
- Regional processing endpoints carry a 10% uplift.
1M tokens · USD
- Input
- $0.75
- Cached input
- $0.075
- Output
- $4.50
Example: 10K input + 2K output
- Input cost
- $0.0075
- Output cost
- $0.0090
- Estimated total
- $0.0165
04 / Context
A 400K Context Window for Focused Agent Work
GPT-5.4 Mini provides a 400,000-token context window and supports up to 128,000 output tokens, giving coding and agent workflows enough room for substantial working context without matching the 1.05M window of full GPT-5.4.
Large Enough for Many Repository and Tool Workflows
A 400K context window can hold a meaningful amount of code, documentation, tool definitions, conversation state, and retrieved information.
For coding assistants, this can support several source files, issue context, repository instructions, and tool results in one request. For subagents, it gives each worker enough context to understand its delegated task without requiring the orchestrator to send the entire project history.
That boundary is useful architecturally. A Mini subagent should receive the context required for its job, not every piece of information available to the parent agent.
OpenAI also lists compaction support for GPT-5.4 Mini. In longer workflows, compaction can help preserve relevant state while reducing the amount of prior interaction that must stay verbatim in the active context.
- 400,000-token context window.
- Maximum output of 128,000 tokens.
- Suitable for substantial coding and subagent working sets.
- Smaller than full GPT-5.4's 1.05M context window.
- Supports compaction for longer-running workflows.
Context window
400,000
Max output
128,000
A focused subagent context can include repository instructions, selected files, task requirements, tool definitions, retrieved information, and intermediate results without carrying the entire parent workflow.
05 / Reasoning
Reasoning from None to XHigh
GPT-5.4 Mini supports none, low, medium, high, and xhigh reasoning effort, with none as the default.
Keep Fast Tasks Fast, Increase Effort Selectively
The default none setting fits Mini's execution-oriented role. Straightforward edits, navigation, extraction, or tool-driven steps can begin without additional reasoning overhead.
More difficult coding or multimodal tasks can move to low, medium, high, or xhigh when evaluation shows a meaningful quality improvement.
This is particularly useful in subagent architectures because different workers do not need identical reasoning budgets. A repository search agent can run at low effort while a debugging agent receives a higher setting.
OpenAI's prompting guidance also matters here: GPT-5.4 Mini is more literal than larger models and less likely to infer missing steps. Clear execution order, explicit decision rules, and precise output expectations can improve reliability without automatically increasing reasoning effort.
Fast delegated task
Use none or low for repository search, targeted edits, classification, and clearly specified tool steps.
Harder coding task
Increase reasoning when debugging, planning, multimodal interpretation, or tool selection becomes less predictable.
06 / Capabilities
A Full Agent Toolset in a Mini Model
GPT-5.4 Mini supports streaming, function calling, structured outputs, web and file search, code execution, computer use, MCP, Tool Search, Skills, Apply Patch, and image input.
Small Model, Broad Execution Surface
The supported toolset makes GPT-5.4 Mini more than a fast text generator.
For coding workflows, Code Interpreter, Hosted Shell, and Apply Patch support executable analysis and code changes. For grounded tasks, Web Search and File Search can retrieve external information. Computer Use enables direct interaction with software interfaces.
Tool Search is particularly relevant in agent systems with many available actions. Instead of loading every tool definition into every prompt, an application can allow the model to discover relevant tools as needed.
MCP and Skills extend the integration surface further, while function calling and structured outputs provide application-level control over how results are returned and executed.
- Streaming responses.
- Function calling and structured outputs.
- Web Search and File Search.
- Image Generation and Code Interpreter.
- Hosted Shell and Apply Patch.
- Skills and Computer Use.
- MCP and Tool Search.
- Fine-tuning is currently not supported.
- Supported
Coding execution
Use Code Interpreter, Hosted Shell, and Apply Patch for technical workflows.
- Supported
Computer use
Interpret visual application state and interact with graphical interfaces.
- Supported
Tool Search
Discover relevant tools dynamically in large agent tool ecosystems.
- Supported
Structured integration
Use function calling and structured outputs for predictable application behavior.
- Supported
Search and files
Ground tasks with Web Search and File Search.
- Supported
Extensible agents
Use MCP, Skills, image generation, and other supported Responses API tools.
07 / Evaluation
Strengths and Limitations
GPT-5.4 Mini offers a strong performance-to-latency profile for coding, subagents, and computer use, but it works best when tasks are clearly structured and the workflow does not require flagship-level ambiguity handling.
Strengths
Strong coding performance for a mini model
OpenAI positions GPT-5.4 Mini as its strongest mini model for coding and reports substantial gains over GPT-5 Mini.
Excellent subagent fit
Its speed, price, and capability make it suitable for parallel delegated tasks inside larger agent systems.
Computer-use capability
The model can interpret screenshots and interact with software interfaces in responsive automation workflows.
Broad tool support
Search, shell, code execution, patching, Skills, MCP, Tool Search, and structured integration support production agents.
What to consider
Smaller context than flagship GPT-5.4
The 400K window is substantial but below the 1.05M context available on full GPT-5.4.
More literal prompting behavior
OpenAI notes that Mini is less likely than larger models to infer missing workflow steps or resolve ambiguity implicitly.
Not the strongest model for every hard task
Complex planning, ambiguous professional work, or final high-stakes evaluation may still benefit from a larger model.
Fine-tuning is unavailable
The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.4 Mini.
Evaluate fast agent execution
Test GPT-5.4 Mini on your real coding and subagent tasks
Run representative coding loops, tool calls, computer-use tasks, and delegated subagent work, then inspect response quality, latency, token usage, cost, and context consumption in EidoStack.
Start FreeConnect your own OpenAI API key and evaluate GPT-5.4 Mini with the same prompts, tools, reasoning settings, and context structure your production workflow will use.
Common Questions
What is GPT-5.4 Mini?
GPT-5.4 Mini is OpenAI's strongest mini model for coding, computer use, and subagents. It brings many GPT-5.4-class capabilities to a faster and more efficient model designed for high-volume workflows.
How much does GPT-5.4 Mini cost?
Standard pricing is $0.75 per 1M input tokens, $0.075 per 1M cached input tokens, and $4.50 per 1M output tokens. Tool-specific charges may apply separately.
What is the context window of GPT-5.4 Mini?
GPT-5.4 Mini has a 400,000-token context window and supports up to 128,000 output tokens.
What reasoning levels does GPT-5.4 Mini support?
GPT-5.4 Mini supports none, low, medium, high, and xhigh reasoning effort. None is the default.
What is the knowledge cutoff for GPT-5.4 Mini?
OpenAI lists August 31, 2025 as the knowledge cutoff for GPT-5.4 Mini.
What workloads are a good fit for GPT-5.4 Mini?
GPT-5.4 Mini is well suited to coding assistants, repository navigation, targeted edits, debugging loops, parallel subagents, computer-use automation, and high-volume tool-calling workflows.
Does GPT-5.4 Mini support computer use?
Yes. Computer Use is supported, and OpenAI specifically highlights GPT-5.4 Mini for interface-control workflows that need to interpret screenshots and act quickly.
What tools does GPT-5.4 Mini support?
OpenAI lists Web Search, File Search, Image Generation, Code Interpreter, Hosted Shell, Apply Patch, Skills, Computer Use, MCP, and Tool Search as supported Responses API tools.
Does GPT-5.4 Mini support structured outputs?
Yes. GPT-5.4 Mini supports structured outputs as well as function calling and streaming.
Why is GPT-5.4 Mini useful for subagents?
Its combination of speed, lower cost, coding capability, tools, and reasoning controls makes it suitable for delegated tasks that can run in parallel under a larger coordinating model.
How should prompts be written for GPT-5.4 Mini?
OpenAI recommends explicit prompts with critical rules first, clear execution order, decision rules, and separate instructions for taking an action versus reporting the action. Mini is more literal and less likely than larger models to infer missing steps.
Can GPT-5.4 Mini be fine-tuned?
No. The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.4 Mini.
Model information
Last updated
The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.4 Mini and OpenAI's GPT-5.4 Mini release guidance. Provider pricing, model limits, tools, processing options, and availability may change, so production assumptions should be checked against the latest provider documentation.