OpenAI model
GPT-5.6 Luna
A fast and cost-efficient model in the GPT-5.6 family for applications where scale, per-request cost, and large-context processing matter.
- Context window
- 1.05M
- tokens
- Max output
- 128K
- tokens
- Input
- $0.20
- per 1M tokens
- Cached input
- $0.02
- per 1M tokens
- Output
- $1.20
- per 1M tokens
01 / Overview
What GPT-5.6 Luna Is
Luna is an OpenAI model for scenarios where cost and scale matter as much as response quality. It occupies the most cost-efficient tier within the GPT-5.6 family.
More Than Just a Cheap Model
Low cost is Luna's main practical advantage, but the model is not limited to the simplest operations. It supports reasoning, image input, function calling, structured outputs, and tools available through the Responses API.
That makes it a candidate not only for classification or data extraction, but also for production workflows where the model must follow instructions, return structured results, and interact with external tools.
- Supports large volumes of repeated API calls.
- Provides a large input context of up to 1.05M tokens.
- Lets you adjust reasoning effort to match task complexity.
- Provider
- OpenAI
- Family
- GPT-5.6
- Status
- Active
- Knowledge cutoff
- Feb 16, 2026
- Input modalities
- Text, Image
- Output modality
- Text
- Default reasoning
- Medium
02 / Use cases
When It Makes Sense to Choose Luna
Luna is especially useful when an AI call needs to be good enough for production but inexpensive enough to run at scale. This matters for products with a constant flow of model requests.
Focus on the Cost of a Useful Result
In production, token price alone is not enough. The model must reliably complete the task at the quality level your application requires. If a cheaper model causes more retries or requires additional validation, the real savings may be smaller than expected.
That is why Luna should be tested on a representative set of real product requests. Measure response quality, output length, token usage, and the actual cost of one useful result.
- Well suited to cost-sensitive production workloads.
- Useful as a lower-cost step inside a multi-stage AI workflow.
- Especially valuable when monthly inference volume is high.
- 01
High-volume requests
Classification, routing, extraction, and background AI operations.
- 02
Long documents
Process large input contexts without aggressively trimming the source material.
- 03
Structured workflows
Scenarios where the application expects a predictable structured result.
- 04
Tool-based automation
Sequences that use function calling and tools from the Responses API.
03 / Pricing
GPT-5.6 Luna Pricing
Low API cost is one of the main reasons to consider Luna for production. It is important to evaluate not only the price per million tokens, but also the structure of a typical request in your application.
Your Context Determines the Real Cost
Products with long system prompts, large conversation histories, or retrieved documents can consume far more input tokens than a short test prompt suggests.
Repeated context can be cheaper when cached input is available. For that reason, real cost should be estimated across several representative scenarios: a short request, an average session, and a context-heavy production request.
- $0.20 per 1M standard input tokens.
- $0.02 per 1M cached input tokens.
- $1.20 per 1M output tokens.
1M tokens · USD
- Input
- $0.20
- Cached input
- $0.02
- Output
- $1.20
Example: 10K input + 2K output
- Input cost
- $0.0020
- Output cost
- $0.0024
- Estimated total
- $0.0044
04 / Context
A Context Window of Up to 1.05 Million Tokens
The large context window allows the model to receive long documents, extended conversation history, detailed instructions, and a substantial amount of retrieved context.
Large Context Should Be Used Deliberately
The maximum window size is a technical limit, not a recommendation to fill it completely. The more data you send with every request, the higher the cost, and the harder it can become for the model to distinguish essential instructions from secondary information.
For production systems, it is useful to measure actual usage: how much space is taken by the system prompt, conversation history, retrieved documents, and the user's new request. This shows how efficiently the application uses the available context.
- Up to 1.05M tokens of total context.
- Up to 128K tokens of maximum output.
- Suitable for document-heavy and long-context workloads.
Context window
1,050,000
Max output
128,000
Context includes system instructions, conversation history, user input, and other data sent to the model in a specific request.
05 / Reasoning
Adjustable Reasoning Effort
Luna lets you choose how much reasoning the model should use based on task complexity instead of applying the same compute level to every request.
Maximum Reasoning Is Not Always Necessary
For simple high-volume operations, a high reasoning effort can be unnecessary. For more complex tasks, such as checking multiple conditions or solving a multi-step problem, increasing effort may produce a more reliable result.
The correct setting depends on the workload. It is useful to test the same task set at several effort levels and evaluate whether the change in quality justifies the change in latency and token consumption.
Simple workload
Classification, extraction, or short transformations usually do not require maximum reasoning effort.
Complex workload
More complex instructions and decisions involving multiple conditions should be tested with higher effort.
06 / Capabilities
A Modern Set of API Capabilities
Luna supports OpenAI API capabilities needed not only for text chat, but also for structured automation and tool-based workflows.
Designed for Application Integration
Function calling allows the model to select functions predefined by the application. Structured outputs are useful when free-form text is inconvenient and the result must match an expected structure.
Responses API tools extend the model beyond a simple prompt-to-response interaction. The model can participate in longer action sequences and work with external information sources.
- Text and image input.
- Streaming, function calling, and structured outputs.
- Web search, file search, and other Responses API tools.
- Supported
Text
Text input and output.
- Supported
Image input
Images can be analyzed as input.
- Supported
Streaming
Receive output progressively.
- Supported
Function calling
Invoke application-defined tools.
- Supported
Structured outputs
Predictable machine-readable results.
- Supported
Responses tools
Web search, file search, and other tools.
07 / Evaluation
Strengths and Limitations
Luna should be evaluated as a cost-efficient production model, not as a universal replacement for larger models across every possible task.
Strengths
Low cost
Useful for scaling a large number of API operations.
Large context
A 1.05M-token window provides substantial capacity for documents and long conversation history.
Modern API
Reasoning, tools, function calling, and structured output are available in one model.
Flexible reasoning
Reasoning effort can be adjusted to match the complexity and economics of a specific workload.
What to consider
Not the highest-capability tier
Luna is optimized for cost and speed; the most difficult workloads may require a stronger model.
Long-context pricing
Very large input prompts cost more and should be used deliberately.
No fine-tuning
If fine-tuning is a hard requirement, a different model path is needed.
Real workload testing is still required
Technical specifications do not replace evaluation on your own prompts and production data.
Make the decision with evidence
Test GPT-5.6 Luna on your own prompts
Use your real system prompt and workload, then inspect the response, token usage, cost, and actual context-window consumption in one workspace.
Start FreeConnect your own provider API key and evaluate the model under the same conditions your application will use.
Common Questions
What is GPT-5.6 Luna?
GPT-5.6 Luna is the cost-efficient tier of the GPT-5.6 family, designed for high-volume and cost-sensitive workloads. It supports reasoning, a large context window, and modern API capabilities.
How much does GPT-5.6 Luna cost?
Standard pricing is $0.20 per 1M input tokens, $0.02 per 1M cached input tokens, and $1.20 per 1M output tokens.
What is the context window of GPT-5.6 Luna?
The context window is 1,050,000 tokens, with a maximum output of 128,000 tokens.
Does GPT-5.6 Luna support reasoning?
Yes. Available reasoning-effort levels are none, low, medium, high, xhigh, and max. Medium is the default level.
Can GPT-5.6 Luna accept images?
Yes. The model supports image input and text output, which allows it to analyze visual content.
What workloads are best suited to GPT-5.6 Luna?
Luna is particularly suitable for high-volume, cost-sensitive operations such as extraction, classification, document processing, structured workflows, and tool-based automation.
Model information
Last updated
The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.6 Luna. For production use, pricing, limits, and capabilities should be checked periodically because providers may update model parameters.