OpenAI lowest-cost GPT-5 tier
GPT-5 Nano
The fastest and cheapest original GPT-5 model, designed for summarization, classification, and high-volume workloads where tiny per-request costs matter.
- Context window
- 400K
- tokens
- Max output
- 128K
- tokens
- Input
- $0.05
- per 1M tokens
- Cached input
- $0.005
- per 1M tokens
- Output
- $0.40
- per 1M tokens
01 / Overview
What GPT-5 Nano Is
GPT-5 Nano is the fastest and most cost-efficient model in the original GPT-5 family, built for simple, repeatable AI work such as summarization and classification where throughput and unit economics dominate model selection.
The GPT-5 Tier Built Around Cost per Operation
Nano models solve a different engineering problem from flagship models.
A flagship is selected when the application needs the strongest available reasoning. A Nano model is selected when a workload happens constantly and each individual operation is narrow enough that maximum intelligence would be unnecessary.
GPT-5 Nano was OpenAI's lowest-cost GPT-5 tier. Its official positioning emphasizes speed and price, with summarization and classification named as representative workloads.
That makes the model particularly relevant to systems processing large streams of messages, records, documents, events, or retrieved text. At sufficient volume, a fraction of a cent per call becomes an infrastructure-level cost decision.
The model still provides a 400K context window, reasoning token support, image input, function calling, and structured outputs. "Nano" therefore describes its efficiency tier rather than a minimal API surface.
GPT-5 Nano is now deprecated. OpenAI recommends GPT-5.6 Luna for most new speed- and cost-sensitive workloads.
- Fastest and cheapest original GPT-5 model.
- Designed for summarization and classification.
- Useful when request volume magnifies small unit costs.
- 400K context and 128K maximum output.
- Deprecated with GPT-5.6 Luna as the recommended migration target.
- Provider
- OpenAI
- Family
- GPT-5
- Tier
- Nano
- Status
- Deprecated
- Knowledge cutoff
- May 31, 2024
- Input modalities
- Text, Image
- Output modality
- Text
02 / Use cases
Where GPT-5 Nano Fit Best
GPT-5 Nano fits narrow tasks that are easy to define, inexpensive to verify, and executed at enough volume that inference cost and throughput become more important than frontier-level reasoning.
Repetition Is Where Nano Economics Become Valuable
The clearest use case is a high-frequency operation with a stable contract.
A support system may classify every incoming message. A content pipeline may summarize thousands of records. A search system may normalize or label retrieved items. An application may need a lightweight decision before routing work to a larger model.
In each case, the application can define the expected output precisely and validate it cheaply.
GPT-5 Nano can also act as the first stage of a model cascade. Straightforward requests stay on the inexpensive tier while uncertain or difficult cases escalate to a stronger model.
This architecture matters more than choosing the cheapest token price in isolation. The real objective is to minimize the cost of a successful workflow without creating excessive retries, false classifications, or unnecessary escalations.
- 01
Classification
Assign categories, labels, priorities, or routing decisions to large streams of clearly defined inputs.
- 02
Summarization
Compress messages, records, retrieved passages, or documents when the expected summary format is explicit.
- 03
Structured extraction
Convert repeated input patterns into schema-constrained fields for downstream systems.
- 04
First-pass routing
Handle simple requests cheaply and escalate uncertain or complex cases to a more capable model.
03 / Pricing
GPT-5 Nano Pricing
GPT-5 Nano costs $0.05 per million input tokens, $0.005 per million cached input tokens, and $0.40 per million output tokens, making it the cheapest member of the original GPT-5 family.
Optimize Millions of Small Decisions
Nano pricing becomes most meaningful when expressed per workload rather than per token.
At $0.05 per million input tokens, ten thousand input tokens cost only $0.0005. Two thousand output tokens at $0.40 per million add $0.0008, producing a $0.0013 example request before other applicable charges.
That is small for one call. At millions of calls, it becomes a material operating cost.
Cached input is priced at $0.005 per million tokens, one tenth of standard input. Stable prompt prefixes and repeated context can therefore improve economics further when caching applies.
Output is eight times the standard input token price, so concise schemas, labels, enums, and short summaries are particularly well aligned with Nano economics.
- $0.05 per 1M input tokens.
- $0.005 per 1M cached input tokens.
- $0.40 per 1M output tokens.
- Cached input costs 10% of uncached input.
- Short structured outputs help preserve the model's cost advantage.
1M tokens · USD
- Input
- $0.05
- Cached input
- $0.005
- Output
- $0.40
Example: 10K input + 2K output
- Input cost
- $0.0005
- Output cost
- $0.0008
- Estimated total
- $0.0013
04 / Context
400K Context Without Flagship Token Pricing
GPT-5 Nano provides a 400,000-token context window and up to 128,000 output tokens, allowing an efficiency-tier model to process substantial source material when a workload genuinely requires it.
Context Capacity Is Useful, but Filtering Still Wins
A large context window can support long documents, retrieved passages, conversation history, instructions, and batches of related records.
For a Nano workload, however, the strongest economic pattern is usually not to fill the window.
Classification and extraction often improve when the model receives only the evidence needed for the decision. Summarization may require longer source material, but irrelevant context still consumes tokens and can dilute the signal.
High-volume systems should therefore treat context selection as part of inference optimization. Retrieval, chunking, filtering, and compact prompt design can reduce both cost and noise.
The 128K maximum output is far larger than most Nano use cases require. For predictable economics, constrain generated length to the task's actual contract.
Context window
400,000
Max output
128,000
A 400K context window provides capacity for large inputs, but high-volume Nano workloads usually benefit from aggressive context filtering and short outputs.
05 / Reasoning
Reasoning Support for Lightweight Decisions
GPT-5 Nano supports reasoning tokens, extending the cheapest GPT-5 tier beyond direct completion into bounded decisions that require interpretation or a small amount of multi-step inference.
Use Reasoning Where It Improves Task Accuracy
Reasoning support does not change Nano's core role.
The model remains optimized around speed and cost, so reasoning should serve a measurable task requirement rather than turn every small operation into an open-ended deliberation.
A classification rule may require reconciling several fields. An extraction task may need to infer which value corresponds to a schema field. A routing step may need to apply multiple constraints before selecting a destination.
These are reasonable places to evaluate reasoning.
The important metric is task-level value. If additional reasoning improves accuracy enough to reduce retries or expensive downstream escalation, it can improve total workflow economics even when it consumes more tokens.
Rule-based classification
Apply several explicit criteria before returning a compact category or structured decision.
Escalation gate
Use a bounded reasoning step to decide whether a request can stay on the low-cost path or needs a stronger model.
06 / Capabilities
Structured Outputs and Function Calling at Nano Cost
GPT-5 Nano supports streaming, function calling, structured outputs, text and image input, and reasoning tokens, giving low-cost inference the controls needed for production automation.
Machine-Readable Output Is Central to Nano Workloads
Many Nano use cases are application operations rather than conversations.
A classifier should return a known category. An extractor should populate known fields. A router should select a known destination. A summarizer may need a predictable object rather than free-form prose.
Structured outputs help enforce those contracts.
Function calling allows the model to participate in tool-driven flows, while image input extends classification and extraction scenarios beyond text-only data.
Streaming is supported, although many Nano workloads produce short structured results where streaming is less important than throughput.
OpenAI lists fine-tuning and predicted outputs as unsupported for GPT-5 Nano.
- Supported
Structured outputs
Constrain classifications, extracted fields, routing decisions, and other application results to a defined schema.
- Supported
Function calling
Connect inexpensive inference to tools and application-defined actions.
- Supported
Streaming
Stream text when incremental delivery is useful for an interactive workload.
- Supported
Image input
Process images together with text while producing text output.
- Not listed
Fine-tuning
The official model card lists fine-tuning as unsupported.
- Not listed
Predicted outputs
Predicted outputs are not supported.
07 / Evaluation
Strengths and Limitations
GPT-5 Nano's defining advantage is extreme unit-cost efficiency for narrow, repeatable work; its limitations appear when tasks become ambiguous, capability-intensive, or valuable enough that a stronger model's higher success rate outweighs Nano's lower token price.
Where GPT-5 Nano stands out
Lowest GPT-5 token price
$0.05 input and $0.40 output per million tokens make GPT-5 Nano the cheapest original GPT-5 tier.
High-volume economics
The model is designed for workloads where tiny per-call savings compound across large numbers of operations.
Strong fit for simple repeatable tasks
OpenAI specifically highlights summarization and classification as representative GPT-5 Nano use cases.
Large context window
A 400K context window allows substantial inputs without moving to a more expensive GPT-5 tier solely for context capacity.
Production output controls
Structured outputs and function calling make Nano useful for automation, routing, and machine-readable pipelines.
What to consider
Deprecated
OpenAI marks GPT-5 Nano as deprecated and recommends GPT-5.6 Luna for most new speed- and cost-sensitive workloads.
Not intended for maximum capability
Complex, ambiguous, or high-value problems may achieve better workflow economics on a stronger model despite its higher token price.
Older knowledge cutoff
The official model card lists May 31, 2024 as the knowledge cutoff.
No fine-tuning
Fine-tuning is not supported for GPT-5 Nano.
Snapshot shutdown
The gpt-5-nano-2025-08-07 snapshot is scheduled for shutdown on December 11, 2026, with GPT-5.6 Luna as the recommended replacement.
Benchmark cost at production scale
Measure GPT-5 Nano against its replacement
Replay representative classification, summarization, extraction, and routing workloads across GPT-5 Nano and newer models, then compare quality, token usage, and effective cost in EidoStack.
Start FreeOpenAI recommends GPT-5.6 Luna for most new speed- and cost-sensitive workloads.
Common Questions
What is GPT-5 Nano?
GPT-5 Nano is the fastest and most cost-efficient version of the original GPT-5 family. OpenAI highlights summarization and classification as representative workloads.
How much does GPT-5 Nano cost?
GPT-5 Nano costs $0.05 per 1M input tokens, $0.005 per 1M cached input tokens, and $0.40 per 1M output tokens.
Is GPT-5 Nano cheaper than GPT-5 Mini?
Yes. GPT-5 Nano is the lowest-cost tier in the original GPT-5 family and is intended for simpler, more cost-sensitive workloads.
What is the GPT-5 Nano context window?
GPT-5 Nano has a 400,000-token context window.
What is the maximum output of GPT-5 Nano?
GPT-5 Nano supports up to 128,000 output tokens.
What is the GPT-5 Nano knowledge cutoff?
The official model card lists May 31, 2024 as the knowledge cutoff.
Does GPT-5 Nano support reasoning?
Yes. OpenAI lists reasoning token support for GPT-5 Nano.
Does GPT-5 Nano support images?
Yes. GPT-5 Nano accepts text and image input and produces text output.
Does GPT-5 Nano support structured outputs?
Yes. Structured outputs are supported and are useful for classification, extraction, and other schema-driven workloads.
Does GPT-5 Nano support function calling?
Yes. The official model card lists function calling as supported.
Can GPT-5 Nano be fine-tuned?
No. Fine-tuning is not supported.
Is GPT-5 Nano deprecated?
Yes. OpenAI currently marks GPT-5 Nano as deprecated.
When will gpt-5-nano-2025-08-07 shut down?
OpenAI's deprecation schedule lists December 11, 2026 as the API shutdown date for gpt-5-nano-2025-08-07.
What replaces GPT-5 Nano?
OpenAI recommends GPT-5.6 Luna for most new speed- and cost-sensitive workloads and specifically lists it as the replacement for gpt-5-nano-2025-08-07.
When should GPT-5 Nano be used instead of a larger model?
GPT-5 Nano is most suitable for narrow, repeatable, high-volume tasks where success criteria are explicit and the cost advantage is more valuable than maximum reasoning capability.
Model information
Last updated
The specifications, pricing, capabilities, positioning, and lifecycle information on this page are based on OpenAI's official GPT-5 Nano model documentation, model guidance, and API deprecation schedule.