OpenAI model

GPT-6 Luna

OpenAI's most efficient GPT-6 model for focused, high-volume tasks where throughput, predictable cost, and modern API capabilities matter.

Context window
1.05M
tokens
Max output
128K
tokens
Input
$0.10
per 1M tokens
Cached input
$0.01
per 1M tokens
Output
$0.50
per 1M tokens

01 / Overview

What GPT-6 Luna Is

GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks, designed for applications that need to run large numbers of well-defined AI operations without making inference cost the dominant part of the product economics.

Efficiency Without Dropping Modern API Features

A high-volume model does not have to be limited to basic text completion. GPT-6 Luna supports configurable reasoning, image input, streaming, function calling, structured outputs, and the same broad Responses API tool categories used by more expensive GPT-6 models.

The important distinction is workload shape. Luna is best aligned with tasks that can be clearly scoped and evaluated: extract fields from documents, classify incoming content, route requests, transform data, summarize bounded inputs, produce structured results, or execute predictable tool-driven steps.

This makes the model particularly relevant when an application performs the same type of AI operation thousands or millions of times. At that scale, a small difference in cost per successful task can materially change monthly inference spend.

  • OpenAI's most efficient model for focused, high-volume tasks.
  • Standard input pricing of $0.10 per million tokens.
  • Supports a 1.05M-token context window despite its efficiency positioning.
  • Includes adjustable reasoning from none through max.
  • Supports vision input, structured outputs, function calling, and Responses API tools.
Model profile
Provider
OpenAI
Family
GPT-6
Positioning
Focused, high-volume
Knowledge cutoff
May 18, 2026
Input modalities
Text, Image
Output modality
Text
Default reasoning
Medium

02 / Use cases

Where GPT-6 Luna Fits Best

GPT-6 Luna is most compelling when a workload is frequent, bounded, and measurable: the application knows what a successful result looks like and needs to produce that result repeatedly at low cost.

Optimize Repeated Work, Not Just Individual Prompts

High-volume AI systems behave differently from occasional interactive use. A prompt that costs fractions of a cent can become a significant expense when it runs continuously across user activity, background jobs, ingestion pipelines, or automated workflows.

Luna's low token pricing makes it a natural candidate for those repeated operations. Examples include extracting structured data from incoming documents, categorizing support requests, routing work to downstream systems, rewriting content into a fixed format, summarizing records, or producing machine-readable outputs for application logic.

It can also act as an economical first stage in a model-routing architecture. A focused request can be handled by Luna, while genuinely difficult cases are escalated to a stronger model. The value of this approach should be measured with real failure rates and escalation thresholds rather than assumed from pricing alone.

  • Classification and request routing at scale.
  • Structured extraction from documents or messages.
  • Repetitive transformations and normalization.
  • Background AI jobs and ingestion pipelines.
  • First-stage processing before selective escalation to a stronger model.
High-volume workloads
  1. 01

    Classification and routing

    Categorize requests, detect intent, assign workflows, and route traffic using predictable output schemas.

  2. 02

    Structured extraction

    Turn documents, messages, and records into machine-readable fields for downstream systems.

  3. 03

    Content transformation

    Rewrite, normalize, summarize, or format bounded inputs as part of repeatable application pipelines.

  4. 04

    Background automation

    Run frequent AI operations where per-task cost and aggregate monthly volume both matter.

03 / Pricing

GPT-6 Luna Pricing

GPT-6 Luna uses an efficiency-first price structure: $0.10 per million standard input tokens, $0.01 per million cached input tokens, $0.125 per million cache-write tokens, and $0.50 per million output tokens.

Small Per-Request Costs Matter at Large Volume

At low request volume, the difference between fractions of a cent can look negligible. At production scale, those fractions compound across millions of requests, repeated context, background processing, and generated output.

Cached input is priced at ten percent of the uncached input rate. That can be important for applications that repeatedly send the same system instructions, schema definitions, policies, or other stable prompt prefixes.

OpenAI also offers different processing tiers. Batch and Flex are priced at 50% of Standard rates, while Fast mode costs 2× the applicable Standard rates. The useful configuration therefore depends not only on model choice but also on whether the workload prioritizes latency, asynchronous throughput, or minimum cost.

  • $0.10 per 1M standard input tokens.
  • $0.01 per 1M cached input tokens.
  • $0.125 per 1M cache-write tokens.
  • $0.50 per 1M output tokens.
  • Tool-specific operations can introduce separate per-call charges.
Token pricing

1M tokens · USD

Input
$0.10
Cached input
$0.01
Cache writes
$0.125
Output
$0.50

Example: 10K input + 2K output

Input cost
$0.0010
Output cost
$0.0010
Estimated total
$0.0020

04 / Context

A 1.05M-Token Window in an Efficiency-Focused Model

GPT-6 Luna combines low token pricing with a 1,050,000-token context window and up to 128,000 output tokens, giving high-volume applications room to process substantial inputs when a task genuinely requires them.

Large Context Should Remain an Exception When Efficiency Is the Goal

Luna can technically accept very large working sets, but an efficiency-focused architecture should still minimize unnecessary context.

For a document pipeline, the model may need an entire contract, report, transcript, or collection of records. For an application workflow, it may need system instructions, retrieved data, prior state, and the current request. The large context window provides flexibility when those inputs are necessary.

However, sending a million tokens simply because the model can accept them works against Luna's cost-efficiency advantage. Extra context increases spend, can introduce irrelevant information, and crosses into higher long-context pricing above 272K input tokens.

The better pattern is to measure which parts of the context actually improve task success, keep stable prefixes cacheable, and reserve very large requests for workloads that truly need them.

  • 1,050,000-token context window.
  • Maximum output of 128,000 tokens.
  • Suitable for large documents and context-heavy extraction when required.
  • Cached stable context can materially reduce repeated input cost.
  • Requests above 272K input tokens use higher pricing multipliers.
Context capacity

Context window

1,050,000

Max output

128,000

Input contextOutput limit

Context may include system instructions, retrieved records, documents, conversation state, tool results, and the current request. Efficient workloads should send only what improves the result.

05 / Reasoning

Reasoning Controls from None to Max

GPT-6 Luna supports none, low, medium, high, xhigh, and max reasoning effort, with medium as the default.

Match Reasoning Spend to Task Difficulty

A model optimized for volume benefits from differentiated reasoning settings. There is little value in applying maximum reasoning to a deterministic extraction task if a lower setting already reaches the required accuracy.

The none setting is particularly relevant for simple, latency-sensitive, or highly structured workloads. More complex requests can progressively use low, medium, or higher effort when evaluation shows a measurable improvement.

API behavior also matters. OpenAI recommends the Responses API for built-in tools and function calling. Chat Completions supports function calling with GPT-6 Luna only when reasoning_effort is set to none.

For production systems, reasoning effort should therefore be part of the workload configuration: evaluate quality, latency, token usage, and tool requirements together rather than choosing one global setting.

Reasoning effort
nonelowmedium · defaulthighxhighmax

Focused high-volume task

Test none or low reasoning for extraction, classification, formatting, and other tightly scoped operations.

More complex workflow

Increase reasoning only when evaluation shows that additional effort materially improves task success.

06 / Capabilities

Production Features Beyond Simple Text Generation

GPT-6 Luna combines efficiency-oriented pricing with image input, streaming, function calling, structured outputs, and a broad Responses API toolset for building automated application workflows.

Low Cost Does Not Mean a Minimal Integration Surface

Structured outputs are valuable for high-volume automation because they let downstream systems receive predictable data instead of parsing free-form prose. Function calling can connect model decisions to application actions, while search and file tools allow the model to retrieve information when a focused task requires external context.

OpenAI lists web search, file search, image generation, code interpreter, hosted shell, Apply Patch, Skills, computer use, MCP, and Tool Search as supported Responses API tools for GPT-6 Luna.

The model accepts text and image input and generates text output. Direct audio and video modalities are not supported. Fine-tuning is currently listed as unsupported.

  • Text input and output with image input.
  • Streaming responses.
  • Function calling and structured outputs.
  • Web search and file search.
  • Image generation and code interpreter.
  • Hosted shell and Apply Patch.
  • Skills, computer use, MCP, and Tool Search.
  • Fine-tuning is currently not supported.
Supported capabilities
  • Structured automation

    Use structured outputs and function calling for predictable application workflows.

    Supported
  • Text and vision

    Process text and image input while producing text output.

    Supported
  • Search and files

    Use web search and file search when a focused task requires external information.

    Supported
  • Coding tools

    Use code interpreter, hosted shell, and Apply Patch in supported Responses API workflows.

    Supported
  • Computer use

    Interact with graphical environments when automation requires interface-level actions.

    Supported
  • Extensible tools

    Use Skills, MCP, Tool Search, image generation, and other supported Responses API tools.

    Supported

07 / Evaluation

Strengths and Limitations

GPT-6 Luna is strongest when the workload is focused enough to benefit from an efficiency-first model and frequent enough for cost per successful task to matter at scale.

Strengths

  • Very low token cost

    Standard pricing of $0.10 input and $0.50 output per million tokens makes Luna suitable for high-volume production traffic.

  • Large context capacity

    A 1.05M-token window gives an efficiency-focused model substantial flexibility for large documents and application context.

  • Broad reasoning range

    Reasoning can be configured from none through max instead of forcing the same inference effort on every task.

  • Full workflow tooling

    Structured outputs, functions, search, code tools, computer use, MCP, and other Responses API tools enable more than simple text completion.

What to consider

  • Focused-task positioning

    Luna is optimized for efficient, well-scoped work; difficult professional or agentic workloads may justify evaluating Sol or Astra.

  • Scale can hide inefficient prompts

    At high request volume, unnecessary context, verbose output, and retries can erase much of the advantage of a low token rate.

  • Chat Completions tool limitation

    Function calling through Chat Completions is supported only with reasoning_effort set to none; broader tool workflows should use Responses API.

  • Fine-tuning is unavailable

    The current OpenAI model documentation lists fine-tuning as unsupported for GPT-6 Luna.

Measure efficiency on real traffic

Test GPT-6 Luna on your production-style prompts

Evaluate the focused tasks your application runs most often, then inspect response quality, token usage, request cost, reasoning settings, and context consumption in one EidoStack workspace.

Start Free

Connect your own OpenAI API key and evaluate GPT-6 Luna with the same prompts, context, and output requirements your application will use.

Common Questions

What is GPT-6 Luna?

GPT-6 Luna is OpenAI's most efficient GPT-6 model for focused, high-volume tasks. It combines low token pricing with configurable reasoning, a large context window, vision input, structured outputs, function calling, and Responses API tools.

How much does GPT-6 Luna cost?

Standard pricing is $0.10 per 1M input tokens, $0.01 per 1M cached input tokens, $0.125 per 1M cache-write tokens, and $0.50 per 1M output tokens. Tool-specific charges may apply separately.

What is the context window of GPT-6 Luna?

GPT-6 Luna has a 1,050,000-token context window and supports up to 128,000 output tokens.

What reasoning levels does GPT-6 Luna support?

GPT-6 Luna supports none, low, medium, high, xhigh, and max reasoning effort. Medium is the default.

What is the knowledge cutoff for GPT-6 Luna?

OpenAI lists May 18, 2026 as the knowledge cutoff for GPT-6 Luna.

What workloads are a good fit for GPT-6 Luna?

GPT-6 Luna is well suited to high-volume, focused operations such as classification, routing, structured extraction, content transformation, summarization, background processing, and other repeatable tasks with measurable outputs.

Does GPT-6 Luna support image input?

Yes. GPT-6 Luna accepts text and image input and generates text output. Direct audio and video modalities are not supported.

Does GPT-6 Luna support function calling and structured outputs?

Yes. Both are supported. OpenAI recommends the Responses API for built-in tools and function calling. Chat Completions function calling is supported only when reasoning_effort is set to none.

What tools does GPT-6 Luna support?

OpenAI currently lists web search, file search, image generation, code interpreter, hosted shell, Apply Patch, Skills, computer use, MCP, and Tool Search as supported Responses API tools.

Can GPT-6 Luna be fine-tuned?

No. The current OpenAI model documentation lists fine-tuning as unsupported for GPT-6 Luna.

How does GPT-6 Luna reduce cost for high-volume applications?

Its standard token rates are low, cached input costs 10% of uncached input, and Batch or Flex processing can reduce applicable Standard rates by 50%. The actual savings still depend on prompt size, output length, retries, tools, and task success rate.

Model information

Last updated

The specifications and prices on this page are based on the official OpenAI documentation for GPT-6 Luna. Provider pricing, processing tiers, model limits, tools, data-residency options, and availability may change, so production assumptions should be checked against the latest provider documentation.

GPT-6 Luna — Pricing, Context Window & High-Volume AI | EidoStack