OpenAI model
GPT-5.6 Terra
A balanced GPT-5.6 model for production workloads that need stronger intelligence while keeping inference cost under control.
- Context window
- 1.05M
- tokens
- Max output
- 128K
- tokens
- Input
- $2.00
- per 1M tokens
- Cached input
- $0.20
- per 1M tokens
- Output
- $12.00
- per 1M tokens
01 / Overview
What GPT-5.6 Terra Is
GPT-5.6 Terra is the balanced tier of OpenAI's GPT-5.6 family, designed for workloads that need substantial model capability without defaulting to flagship-level pricing for every request.
Built for the Middle of the Production Stack
Terra sits between a cost-first model strategy and the most expensive capability tier. OpenAI describes it as a model that balances intelligence and cost, and notes that it roughly corresponds to the mini tier used in earlier GPT-5 families.
That positioning makes Terra relevant when a workload is too demanding for a purely throughput-oriented model but still needs predictable API economics. It supports configurable reasoning, image input, structured outputs, function calling, and the broader tool ecosystem available through the Responses API.
For product teams, this makes Terra a useful candidate for application features where answer quality matters enough to justify more inference spend, while request volume is still large enough that model cost must remain visible.
- Designed around a balance between model intelligence and inference cost.
- Supports a 1.05M-token context window for context-heavy applications.
- Offers adjustable reasoning effort from
nonethroughmax. - Includes modern tool-calling and structured-output capabilities.
- Provider
- OpenAI
- Family
- GPT-5.6
- Status
- Active
- Knowledge cutoff
- Feb 16, 2026
- Input modalities
- Text, Image
- Output modality
- Text
- Default reasoning
- Medium
02 / Use cases
Where GPT-5.6 Terra Fits Best
Terra is a strong candidate for production features where the workload requires more deliberate reasoning, richer instructions, or tool use, but the economics still need to work at sustained application scale.
A General-Purpose Production Tier
Many AI features do not need the most capable model available, but they also cannot tolerate weak instruction following or overly shallow reasoning. Terra targets that middle ground.
Typical examples include business-process automation, document analysis, structured decision support, application assistants, coding-related workflows, and tool-driven tasks where one request may need to interpret several pieces of context before producing a useful result.
The right fit should still be established empirically. A representative evaluation should include easy, average, and difficult examples from the real workload so that the team can see where Terra remains reliable and where a different model tier may be justified.
- Useful for production workflows with moderate-to-high task complexity.
- Suitable for tool-driven application features and structured automation.
- Practical for long-document and large-context processing.
- Relevant when quality improvements are valuable but flagship cost is difficult to justify across every request.
- 01
Business workflow automation
Process instructions, business rules, retrieved context, and structured application actions.
- 02
Document analysis
Work across long reports, policies, specifications, and other context-heavy source material.
- 03
Application assistants
Power product features that need stronger instruction following and multi-step responses.
- 04
Tool-driven tasks
Use function calling, search, code execution, or other tools as part of a larger workflow.
03 / Pricing
GPT-5.6 Terra Pricing
GPT-5.6 Terra is priced at $2.00 per million input tokens, $0.20 per million cached input tokens, and $12.00 per million output tokens under OpenAI's standard text-token pricing.
Model Cost Is Only Part of Workload Cost
Terra's pricing makes prompt design and output control important. Long context can increase input spend, while verbose responses can make output tokens a significant part of the total request cost.
For applications with repeated system instructions or stable prefixes, cached input can materially change the economics. For workflows that generate long explanations, plans, or structured artifacts, output usage deserves separate attention because output tokens are priced substantially higher than standard input tokens.
A useful production estimate should therefore measure the complete request shape rather than multiplying a single headline rate by monthly request volume.
- $2.00 per 1M standard input tokens.
- $0.20 per 1M cached input tokens.
- $12.00 per 1M output tokens.
- Tool-specific operations may introduce additional charges depending on the tool used.
1M tokens · USD
- Input
- $2.00
- Cached input
- $0.20
- Output
- $12.00
Example: 10K input + 2K output
- Input cost
- $0.0200
- Output cost
- $0.0240
- Estimated total
- $0.0440
04 / Context
1.05M Tokens for Context-Heavy Workloads
GPT-5.6 Terra provides a 1,050,000-token context window and supports up to 128,000 output tokens, giving applications substantial room for documents, history, instructions, and retrieved information.
Capacity Is Not the Same as Context Quality
A large window gives engineers more freedom in how they assemble prompts, but using more context is not automatically better. Irrelevant documents, duplicated instructions, and excessive history can increase cost while making the important parts of the request harder to isolate.
For retrieval-heavy applications, the more useful question is not simply whether the model can fit the available data. It is whether the selected context improves the result enough to justify the additional tokens.
Terra's large window is therefore most valuable when paired with deliberate context management: selecting relevant sources, measuring prompt composition, and tracking how much of the available capacity is actually used in normal production sessions.
- 1,050,000-token context window.
- Maximum output of 128,000 tokens.
- Suitable for long documents, extended sessions, and retrieval-augmented workflows.
- Large-context requests require additional cost discipline above the 272K-input threshold.
Context window
1,050,000
Max output
128,000
The context budget may include system instructions, conversation history, retrieved documents, tool results, and the current user request.
05 / Reasoning
Reasoning Effort from None to Max
GPT-5.6 Terra supports six reasoning-effort levels, allowing developers to tune how much reasoning the model uses instead of treating every request as equally difficult.
Reasoning Is a Workload Parameter
The available levels are none, low, medium, high, xhigh, and max, with medium as the default. This gives teams a practical way to test whether more reasoning improves a particular task enough to justify any additional latency or token usage.
For predictable transformations or straightforward extraction, a lower effort may be sufficient. More involved tasks that require planning, reconciling multiple constraints, or coordinating tool calls may benefit from additional reasoning.
The most useful setting is workload-specific. Rather than selecting the highest level globally, evaluate multiple effort levels against the same test set and compare quality, cost, token usage, and response time.
Routine application task
Extraction, rewriting, or deterministic structured work can be tested first with lower reasoning effort.
Multi-step application task
Planning, constraint-heavy decisions, and tool orchestration may justify testing higher reasoning levels.
06 / Capabilities
Capabilities for Production Applications
Terra combines text and image input with streaming, function calling, structured outputs, and a broad set of Responses API tools, making it suitable for more than conversational interfaces.
From Model Response to Application Workflow
Structured outputs help applications request predictable machine-readable results instead of parsing arbitrary prose. Function calling allows the model to select actions exposed by the application, while Responses API tools extend the workflow to search, files, code execution, shell operations, computer use, MCP, and other supported capabilities.
This matters for production systems because the model can participate in a controlled sequence of application steps rather than operating only as a text generator.
Fine-tuning is not supported for GPT-5.6 Terra according to the current model documentation, so workload adaptation should rely on prompts, context, tools, and application-level orchestration unless OpenAI changes that capability later.
- Text input and output with image input.
- Streaming responses.
- Function calling and structured outputs.
- Responses API tools including web search, file search, code interpreter, hosted shell, computer use, MCP, and additional supported tools.
- Fine-tuning is currently not supported.
- Supported
Text
Text input and text output.
- Supported
Image input
Analyze images supplied as model input.
- Supported
Streaming
Return generated output progressively.
- Supported
Function calling
Connect model decisions to application-defined functions.
- Supported
Structured outputs
Produce responses that conform to a defined structure.
- Supported
Responses tools
Use search, files, code execution, computer use, MCP, and other supported tools.
07 / Evaluation
Strengths and Limitations
GPT-5.6 Terra is most compelling when a product needs a capable general-purpose model with better cost discipline than a flagship-first architecture.
Strengths
Balanced production positioning
Designed specifically for workloads that need to balance intelligence and cost rather than optimize for only one side.
Large context capacity
A 1.05M-token context window supports long documents, retrieved information, and extended application state.
Configurable reasoning
Six reasoning-effort levels make it possible to tune the model for different task classes.
Broad tool support
Function calling and Responses API tools support structured, action-oriented application workflows.
What to consider
Higher cost than throughput-first models
Terra's stronger capability tier comes with higher token pricing, which matters for very large request volumes.
Long prompts can change the economics
Requests above 272K input tokens use higher pricing multipliers, so context size should be monitored carefully.
Fine-tuning is unavailable
The current model documentation lists fine-tuning as unsupported.
Balanced does not mean optimal for every task
Real prompts should be evaluated to determine whether Terra provides enough additional value for its cost on a specific workload.
Evaluate before production
Test GPT-5.6 Terra with your real workload
Run the prompts your application actually uses and inspect response quality, token consumption, request cost, and context usage in one EidoStack workspace.
Start FreeUse your own provider API key and evaluate Terra under the same prompt, context, and reasoning settings you plan to use in production.
Common Questions
What is GPT-5.6 Terra?
GPT-5.6 Terra is an OpenAI model in the GPT-5.6 family designed to balance intelligence and cost. OpenAI describes it as roughly corresponding to the mini model tier used in earlier GPT-5 families.
How much does GPT-5.6 Terra cost?
Standard text-token pricing is $2.00 per 1M input tokens, $0.20 per 1M cached input tokens, and $12.00 per 1M output tokens. Tool-specific charges may apply separately.
What is the context window of GPT-5.6 Terra?
GPT-5.6 Terra has a 1,050,000-token context window and supports up to 128,000 output tokens.
What reasoning levels does GPT-5.6 Terra support?
The model supports none, low, medium, high, xhigh, and max reasoning effort. Medium is the default.
Does GPT-5.6 Terra support image input?
Yes. GPT-5.6 Terra supports text and image input and produces text output. Audio is not supported as a model modality.
Does GPT-5.6 Terra support function calling and structured outputs?
Yes. The model supports both function calling and structured outputs, along with multiple tools through the Responses API.
What types of applications are a good fit for GPT-5.6 Terra?
Terra is a practical candidate for document analysis, business workflow automation, application assistants, structured decision support, coding-related workflows, and tool-driven tasks where both model capability and inference cost matter.
Can GPT-5.6 Terra be fine-tuned?
No. The current OpenAI model documentation lists fine-tuning as unsupported for GPT-5.6 Terra.
Model information
Last updated
The specifications and prices on this page are based on the official OpenAI documentation for GPT-5.6 Terra. Provider pricing, model limits, supported tools, and availability may change, so production assumptions should be checked against the latest provider documentation.