OpenAI model

GPT-5.6 Luna

A fast and cost-efficient model in the GPT-5.6 family for applications where scale, per-request cost, and large-context processing matter.

Context window
1.05M
tokens
Max output
128K
tokens
Input
$0.20
per 1M tokens
Cached input
$0.02
per 1M tokens
Output
$1.20
per 1M tokens

01 / Overview

What GPT-5.6 Luna Is

Luna is an OpenAI model for scenarios where cost and scale matter as much as response quality. It occupies the most cost-efficient tier within the GPT-5.6 family.

More Than Just a Cheap Model

Low cost is Luna's main practical advantage, but the model is not limited to the simplest operations. It supports reasoning, image input, function calling, structured outputs, and tools available through the Responses API.

That makes it a candidate not only for classification or data extraction, but also for production workflows where the model must follow instructions, return structured results, and interact with external tools.

  • Supports large volumes of repeated API calls.
  • Provides a large input context of up to 1.05M tokens.
  • Lets you adjust reasoning effort to match task complexity.
Model profile
Provider
OpenAI
Family
GPT-5.6
Status
Active
Knowledge cutoff
Feb 16, 2026
Input modalities
Text, Image
Output modality
Text
Default reasoning
Medium

02 / Use cases

When It Makes Sense to Choose Luna

Luna is especially useful when an AI call needs to be good enough for production but inexpensive enough to run at scale. This matters for products with a constant flow of model requests.

Focus on the Cost of a Useful Result

In production, token price alone is not enough. The model must reliably complete the task at the quality level your application requires. If a cheaper model causes more retries or requires additional validation, the real savings may be smaller than expected.

That is why Luna should be tested on a representative set of real product requests. Measure response quality, output length, token usage, and the actual cost of one useful result.

  • Well suited to cost-sensitive production workloads.
  • Useful as a lower-cost step inside a multi-stage AI workflow.
  • Especially valuable when monthly inference volume is high.
Typical workloads
  1. 01

    High-volume requests

    Classification, routing, extraction, and background AI operations.

  2. 02

    Long documents

    Process large input contexts without aggressively trimming the source material.

  3. 03

    Structured workflows

    Scenarios where the application expects a predictable structured result.

  4. 04

    Tool-based automation

    Sequences that use function calling and tools from the Responses API.

03 / Pricing

GPT-5.6 Luna Pricing

Low API cost is one of the main reasons to consider Luna for production. It is important to evaluate not only the price per million tokens, but also the structure of a typical request in your application.

Your Context Determines the Real Cost

Products with long system prompts, large conversation histories, or retrieved documents can consume far more input tokens than a short test prompt suggests.

Repeated context can be cheaper when cached input is available. For that reason, real cost should be estimated across several representative scenarios: a short request, an average session, and a context-heavy production request.

  • $0.20 per 1M standard input tokens.
  • $0.02 per 1M cached input tokens.
  • $1.20 per 1M output tokens.
Token pricing

1M tokens · USD

Input
$0.20
Cached input
$0.02
Output
$1.20

Example: 10K input + 2K output

Input cost
$0.0020
Output cost
$0.0024
Estimated total
$0.0044

04 / Context

A Context Window of Up to 1.05 Million Tokens

The large context window allows the model to receive long documents, extended conversation history, detailed instructions, and a substantial amount of retrieved context.

Large Context Should Be Used Deliberately

The maximum window size is a technical limit, not a recommendation to fill it completely. The more data you send with every request, the higher the cost, and the harder it can become for the model to distinguish essential instructions from secondary information.

For production systems, it is useful to measure actual usage: how much space is taken by the system prompt, conversation history, retrieved documents, and the user's new request. This shows how efficiently the application uses the available context.

  • Up to 1.05M tokens of total context.
  • Up to 128K tokens of maximum output.
  • Suitable for document-heavy and long-context workloads.
Context capacity

Context window

1,050,000

Max output

128,000

Input contextOutput limit

Context includes system instructions, conversation history, user input, and other data sent to the model in a specific request.

05 / Reasoning

Adjustable Reasoning Effort

Luna lets you choose how much reasoning the model should use based on task complexity instead of applying the same compute level to every request.

Maximum Reasoning Is Not Always Necessary

For simple high-volume operations, a high reasoning effort can be unnecessary. For more complex tasks, such as checking multiple conditions or solving a multi-step problem, increasing effort may produce a more reliable result.

The correct setting depends on the workload. It is useful to test the same task set at several effort levels and evaluate whether the change in quality justifies the change in latency and token consumption.

Reasoning effort
nonelowmedium · defaulthighxhighmax

Simple workload

Classification, extraction, or short transformations usually do not require maximum reasoning effort.

Complex workload

More complex instructions and decisions involving multiple conditions should be tested with higher effort.

06 / Capabilities

A Modern Set of API Capabilities

Luna supports OpenAI API capabilities needed not only for text chat, but also for structured automation and tool-based workflows.

Designed for Application Integration

Function calling allows the model to select functions predefined by the application. Structured outputs are useful when free-form text is inconvenient and the result must match an expected structure.

Responses API tools extend the model beyond a simple prompt-to-response interaction. The model can participate in longer action sequences and work with external information sources.

  • Text and image input.
  • Streaming, function calling, and structured outputs.
  • Web search, file search, and other Responses API tools.
Supported capabilities
  • Text

    Text input and output.

    Supported
  • Image input

    Images can be analyzed as input.

    Supported
  • Streaming

    Receive output progressively.

    Supported
  • Function calling

    Invoke application-defined tools.

    Supported
  • Structured outputs

    Predictable machine-readable results.

    Supported
  • Responses tools

    Web search, file search, and other tools.

    Supported

07 / Evaluation

Strengths and Limitations

Luna should be evaluated as a cost-efficient production model, not as a universal replacement for larger models across every possible task.

Strengths

  • Low cost

    Useful for scaling a large number of API operations.

  • Large context

    A 1.05M-token window provides substantial capacity for documents and long conversation history.

  • Modern API

    Reasoning, tools, function calling, and structured output are available in one model.

  • Flexible reasoning

    Reasoning effort can be adjusted to match the complexity and economics of a specific workload.

What to consider

  • Not the highest-capability tier

    Luna is optimized for cost and speed; the most difficult workloads may require a stronger model.

  • Long-context pricing

    Very large input prompts cost more and should be used deliberately.

  • No fine-tuning

    If fine-tuning is a hard requirement, a different model path is needed.

  • Real workload testing is still required

    Technical specifications do not replace evaluation on your own prompts and production data.

Make the decision with evidence

Test GPT-5.6 Luna on your own prompts

Use your real system prompt and workload, then inspect the response, token usage, cost, and actual context-window consumption in one workspace.

Start Free

Connect your own provider API key and evaluate the model under the same conditions your application will use.

Common Questions

What is GPT-5.6 Luna?

GPT-5.6 Luna is the cost-efficient tier of the GPT-5.6 family, designed for high-volume and cost-sensitive workloads. It supports reasoning, a large context window, and modern API capabilities.

How much does GPT-5.6 Luna cost?

Standard pricing is $0.20 per 1M input tokens, $0.02 per 1M cached input tokens, and $1.20 per 1M output tokens.

What is the context window of GPT-5.6 Luna?

The context window is 1,050,000 tokens, with a maximum output of 128,000 tokens.

Does GPT-5.6 Luna support reasoning?

Yes. Available reasoning-effort levels are none, low, medium, high, xhigh, and max. Medium is the default level.

Can GPT-5.6 Luna accept images?

Yes. The model supports image input and text output, which allows it to analyze visual content.

What workloads are best suited to GPT-5.6 Luna?

Luna is particularly suitable for high-volume, cost-sensitive operations such as extraction, classification, document processing, structured workflows, and tool-based automation.

Model information

Last updated

The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.6 Luna. For production use, pricing, limits, and capabilities should be checked periodically because providers may update model parameters.

GPT-5.6 Luna — Pricing, Context Window & Capabilities | EidoStack