OpenAI model

GPT-5.6 Terra

A balanced GPT-5.6 model for production workloads that need stronger intelligence while keeping inference cost under control.

Context window
1.05M
tokens
Max output
128K
tokens
Input
$2.00
per 1M tokens
Cached input
$0.20
per 1M tokens
Output
$12.00
per 1M tokens

01 / Overview

What GPT-5.6 Terra Is

GPT-5.6 Terra is the balanced tier of OpenAI's GPT-5.6 family, designed for workloads that need substantial model capability without defaulting to flagship-level pricing for every request.

Built for the Middle of the Production Stack

Terra sits between a cost-first model strategy and the most expensive capability tier. OpenAI describes it as a model that balances intelligence and cost, and notes that it roughly corresponds to the mini tier used in earlier GPT-5 families.

That positioning makes Terra relevant when a workload is too demanding for a purely throughput-oriented model but still needs predictable API economics. It supports configurable reasoning, image input, structured outputs, function calling, and the broader tool ecosystem available through the Responses API.

For product teams, this makes Terra a useful candidate for application features where answer quality matters enough to justify more inference spend, while request volume is still large enough that model cost must remain visible.

  • Designed around a balance between model intelligence and inference cost.
  • Supports a 1.05M-token context window for context-heavy applications.
  • Offers adjustable reasoning effort from none through max.
  • Includes modern tool-calling and structured-output capabilities.
Model profile
Provider
OpenAI
Family
GPT-5.6
Status
Active
Knowledge cutoff
Feb 16, 2026
Input modalities
Text, Image
Output modality
Text
Default reasoning
Medium

02 / Use cases

Where GPT-5.6 Terra Fits Best

Terra is a strong candidate for production features where the workload requires more deliberate reasoning, richer instructions, or tool use, but the economics still need to work at sustained application scale.

A General-Purpose Production Tier

Many AI features do not need the most capable model available, but they also cannot tolerate weak instruction following or overly shallow reasoning. Terra targets that middle ground.

Typical examples include business-process automation, document analysis, structured decision support, application assistants, coding-related workflows, and tool-driven tasks where one request may need to interpret several pieces of context before producing a useful result.

The right fit should still be established empirically. A representative evaluation should include easy, average, and difficult examples from the real workload so that the team can see where Terra remains reliable and where a different model tier may be justified.

  • Useful for production workflows with moderate-to-high task complexity.
  • Suitable for tool-driven application features and structured automation.
  • Practical for long-document and large-context processing.
  • Relevant when quality improvements are valuable but flagship cost is difficult to justify across every request.
Typical workloads
  1. 01

    Business workflow automation

    Process instructions, business rules, retrieved context, and structured application actions.

  2. 02

    Document analysis

    Work across long reports, policies, specifications, and other context-heavy source material.

  3. 03

    Application assistants

    Power product features that need stronger instruction following and multi-step responses.

  4. 04

    Tool-driven tasks

    Use function calling, search, code execution, or other tools as part of a larger workflow.

03 / Pricing

GPT-5.6 Terra Pricing

GPT-5.6 Terra is priced at $2.00 per million input tokens, $0.20 per million cached input tokens, and $12.00 per million output tokens under OpenAI's standard text-token pricing.

Model Cost Is Only Part of Workload Cost

Terra's pricing makes prompt design and output control important. Long context can increase input spend, while verbose responses can make output tokens a significant part of the total request cost.

For applications with repeated system instructions or stable prefixes, cached input can materially change the economics. For workflows that generate long explanations, plans, or structured artifacts, output usage deserves separate attention because output tokens are priced substantially higher than standard input tokens.

A useful production estimate should therefore measure the complete request shape rather than multiplying a single headline rate by monthly request volume.

  • $2.00 per 1M standard input tokens.
  • $0.20 per 1M cached input tokens.
  • $12.00 per 1M output tokens.
  • Tool-specific operations may introduce additional charges depending on the tool used.
Token pricing

1M tokens · USD

Input
$2.00
Cached input
$0.20
Output
$12.00

Example: 10K input + 2K output

Input cost
$0.0200
Output cost
$0.0240
Estimated total
$0.0440

04 / Context

1.05M Tokens for Context-Heavy Workloads

GPT-5.6 Terra provides a 1,050,000-token context window and supports up to 128,000 output tokens, giving applications substantial room for documents, history, instructions, and retrieved information.

Capacity Is Not the Same as Context Quality

A large window gives engineers more freedom in how they assemble prompts, but using more context is not automatically better. Irrelevant documents, duplicated instructions, and excessive history can increase cost while making the important parts of the request harder to isolate.

For retrieval-heavy applications, the more useful question is not simply whether the model can fit the available data. It is whether the selected context improves the result enough to justify the additional tokens.

Terra's large window is therefore most valuable when paired with deliberate context management: selecting relevant sources, measuring prompt composition, and tracking how much of the available capacity is actually used in normal production sessions.

  • 1,050,000-token context window.
  • Maximum output of 128,000 tokens.
  • Suitable for long documents, extended sessions, and retrieval-augmented workflows.
  • Large-context requests require additional cost discipline above the 272K-input threshold.
Context capacity

Context window

1,050,000

Max output

128,000

Input contextOutput limit

The context budget may include system instructions, conversation history, retrieved documents, tool results, and the current user request.

05 / Reasoning

Reasoning Effort from None to Max

GPT-5.6 Terra supports six reasoning-effort levels, allowing developers to tune how much reasoning the model uses instead of treating every request as equally difficult.

Reasoning Is a Workload Parameter

The available levels are none, low, medium, high, xhigh, and max, with medium as the default. This gives teams a practical way to test whether more reasoning improves a particular task enough to justify any additional latency or token usage.

For predictable transformations or straightforward extraction, a lower effort may be sufficient. More involved tasks that require planning, reconciling multiple constraints, or coordinating tool calls may benefit from additional reasoning.

The most useful setting is workload-specific. Rather than selecting the highest level globally, evaluate multiple effort levels against the same test set and compare quality, cost, token usage, and response time.

Reasoning effort
nonelowmedium · defaulthighxhighmax

Routine application task

Extraction, rewriting, or deterministic structured work can be tested first with lower reasoning effort.

Multi-step application task

Planning, constraint-heavy decisions, and tool orchestration may justify testing higher reasoning levels.

06 / Capabilities

Capabilities for Production Applications

Terra combines text and image input with streaming, function calling, structured outputs, and a broad set of Responses API tools, making it suitable for more than conversational interfaces.

From Model Response to Application Workflow

Structured outputs help applications request predictable machine-readable results instead of parsing arbitrary prose. Function calling allows the model to select actions exposed by the application, while Responses API tools extend the workflow to search, files, code execution, shell operations, computer use, MCP, and other supported capabilities.

This matters for production systems because the model can participate in a controlled sequence of application steps rather than operating only as a text generator.

Fine-tuning is not supported for GPT-5.6 Terra according to the current model documentation, so workload adaptation should rely on prompts, context, tools, and application-level orchestration unless OpenAI changes that capability later.

  • Text input and output with image input.
  • Streaming responses.
  • Function calling and structured outputs.
  • Responses API tools including web search, file search, code interpreter, hosted shell, computer use, MCP, and additional supported tools.
  • Fine-tuning is currently not supported.
Supported capabilities
  • Text

    Text input and text output.

    Supported
  • Image input

    Analyze images supplied as model input.

    Supported
  • Streaming

    Return generated output progressively.

    Supported
  • Function calling

    Connect model decisions to application-defined functions.

    Supported
  • Structured outputs

    Produce responses that conform to a defined structure.

    Supported
  • Responses tools

    Use search, files, code execution, computer use, MCP, and other supported tools.

    Supported

07 / Evaluation

Strengths and Limitations

GPT-5.6 Terra is most compelling when a product needs a capable general-purpose model with better cost discipline than a flagship-first architecture.

Strengths

  • Balanced production positioning

    Designed specifically for workloads that need to balance intelligence and cost rather than optimize for only one side.

  • Large context capacity

    A 1.05M-token context window supports long documents, retrieved information, and extended application state.

  • Configurable reasoning

    Six reasoning-effort levels make it possible to tune the model for different task classes.

  • Broad tool support

    Function calling and Responses API tools support structured, action-oriented application workflows.

What to consider

  • Higher cost than throughput-first models

    Terra's stronger capability tier comes with higher token pricing, which matters for very large request volumes.

  • Long prompts can change the economics

    Requests above 272K input tokens use higher pricing multipliers, so context size should be monitored carefully.

  • Fine-tuning is unavailable

    The current model documentation lists fine-tuning as unsupported.

  • Balanced does not mean optimal for every task

    Real prompts should be evaluated to determine whether Terra provides enough additional value for its cost on a specific workload.

Evaluate before production

Test GPT-5.6 Terra with your real workload

Run the prompts your application actually uses and inspect response quality, token consumption, request cost, and context usage in one EidoStack workspace.

Start Free

Use your own provider API key and evaluate Terra under the same prompt, context, and reasoning settings you plan to use in production.

Common Questions

What is GPT-5.6 Terra?

GPT-5.6 Terra is an OpenAI model in the GPT-5.6 family designed to balance intelligence and cost. OpenAI describes it as roughly corresponding to the mini model tier used in earlier GPT-5 families.

How much does GPT-5.6 Terra cost?

Standard text-token pricing is $2.00 per 1M input tokens, $0.20 per 1M cached input tokens, and $12.00 per 1M output tokens. Tool-specific charges may apply separately.

What is the context window of GPT-5.6 Terra?

GPT-5.6 Terra has a 1,050,000-token context window and supports up to 128,000 output tokens.

What reasoning levels does GPT-5.6 Terra support?

The model supports none, low, medium, high, xhigh, and max reasoning effort. Medium is the default.

Does GPT-5.6 Terra support image input?

Yes. GPT-5.6 Terra supports text and image input and produces text output. Audio is not supported as a model modality.

Does GPT-5.6 Terra support function calling and structured outputs?

Yes. The model supports both function calling and structured outputs, along with multiple tools through the Responses API.

What types of applications are a good fit for GPT-5.6 Terra?

Terra is a practical candidate for document analysis, business workflow automation, application assistants, structured decision support, coding-related workflows, and tool-driven tasks where both model capability and inference cost matter.

Can GPT-5.6 Terra be fine-tuned?

No. The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.6 Terra.

Model information

Last updated

The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.6 Terra. Provider pricing, model limits, supported tools, and availability may change, so production assumptions should be checked against the latest provider documentation.

GPT-5.6 Terra — Pricing, Context Window & Reasoning | EidoStack