Anthropic legacy frontier model

Claude Opus 4.6

A transitional Opus model for complex agentic and long-horizon work that introduced adaptive thinking and effort controls while still retaining the deprecated manual extended-thinking API.

Context window
1M
tokens
Max output
128K
tokens
Input
$5.00
per 1M tokens
Cache read
$0.50
per 1M tokens
Output
$25.00
per 1M tokens

01 / Overview

What Claude Opus 4.6 Is

Claude Opus 4.6 is an Anthropic legacy frontier model released for complex agentic tasks and long-horizon work, and it occupies a particularly important point in the evolution of Claude’s reasoning API.

The Opus Generation Where the Reasoning Contract Started to Change

Anthropic released Claude Opus 4.6 on February 5, 2026. It introduced adaptive thinking and generally available effort controls while still accepting the older manual extended-thinking format based on budget_tokens.

That makes Opus 4.6 different from both earlier and later generations. Earlier Claude 4 models relied on manual extended thinking. Opus 4.7 and newer Opus models removed that legacy mode and require adaptive thinking when reasoning is enabled.

Opus 4.6 also helped establish the modern large-context Opus baseline. It now provides a 1,000,000-token context window at standard pricing, supports up to 128,000 standard output tokens, accepts text and images, and supports tool-oriented workflows.

Anthropic lists the model as Active (legacy) and recommends migrating to Claude Opus 5.5 for improved performance.

  • Model ID: claude-opus-4-6.
  • Released February 5, 2026.
  • Active legacy model.
  • 1M context window.
  • 128K standard maximum output.
  • Text and image input with text output.
  • Adaptive thinking supported; legacy extended thinking still accepted but deprecated.
Model profile
Provider
Anthropic
Family
Claude Opus 4.6
Model ID
claude-opus-4-6
Knowledge cutoff
May 2025
Training cutoff
Aug 2025
Input modalities
Text, Image
Lifecycle
Active · legacy

02 / Thinking transition

Claude Opus 4.6 Supports Both Adaptive and Legacy Extended Thinking

Opus 4.6 is the unusual transition point where Anthropic recommends adaptive thinking but still accepts the older thinking: enabled plus budget_tokens interface for backward compatibility.

From Fixed Thinking Budgets to Behavioral Effort

Earlier Claude integrations controlled reasoning depth by explicitly allocating a token budget. A request could set thinking.type to enabled and specify a budget_tokens value.

Opus 4.6 introduced a newer approach: thinking: {"type": "adaptive"} together with the effort parameter. Instead of asking the application to decide how many thinking tokens a task needs, effort acts as a behavioral signal that lets the model decide how much reasoning is appropriate.

The legacy extended-thinking format still works on Opus 4.6, but Anthropic marks it deprecated. Starting with Opus 4.7, the old budget_tokens path returns an error.

Thinking is also off by default on Opus 4.6. If the application does not request thinking, the model runs without it. That differs sharply from newer Claude 5 Opus models, where thinking is on or always on by default.

For teams with older reasoning infrastructure, Opus 4.6 is therefore a useful compatibility baseline: compare the same workload under no thinking, adaptive thinking, and legacy extended thinking before migrating.

Two reasoning APIs
off · default when omittedlowmediumhigh · default effortmax

Adaptive path

Use thinking type adaptive plus effort. This is Anthropic’s recommended reasoning interface for Opus 4.6.

Legacy compatibility path

thinking type enabled with budget_tokens still works on Opus 4.6 but is deprecated and must be removed when moving to Opus 4.7 or later.

03 / Agentic work

Complex Agentic Tasks and Long-Horizon Work

Anthropic launched Claude Opus 4.6 as its most intelligent model at the time for complex agentic tasks and work that unfolds over long sequences of reasoning and tool use.

Judge the Entire Execution Path

Long-horizon agents operate differently from ordinary chat. A coding agent may inspect files, call tools, modify code, run tests, investigate failures, update its plan, and continue for many turns before producing a final answer.

The same pattern appears outside software engineering. A research agent may retrieve evidence repeatedly, a data workflow may call specialized tools, and an enterprise agent may maintain instructions across a long chain of actions.

For these workloads, the useful unit of evaluation is not one response. Measure the complete trajectory:

  • Did the model finish the task?
  • How many tool calls were necessary?
  • Did it repeat work?
  • Did it recover from errors?
  • How much context accumulated?
  • How many tokens were spent on reasoning and output?
  • How often did a human need to intervene?

Opus 4.6’s 1M context window and tool support make it suitable for this style of workload, but later Opus versions improve both model capability and integration behavior. Historical production data from 4.6 is therefore valuable as a regression baseline.

Long-horizon workloads
  1. 01

    Repository-scale coding

    Inspect large codebases, use tools, edit multiple files, run validation, and continue until the implementation is complete.

  2. 02

    Research agents

    Accumulate evidence across many steps while maintaining the original research objective and source context.

  3. 03

    Tool-heavy workflows

    Coordinate repeated function and tool calls where each result influences the next action.

  4. 04

    Long-running enterprise tasks

    Maintain policies, documents, intermediate decisions, and task state across an extended execution path.

04 / Compaction

Opus 4.6 and the Compaction Era

Claude Opus 4.6 launched alongside Anthropic’s compaction API, which introduced server-side context summarization for conversations that need to continue beyond a single raw context window.

A 1M Context Window Still Needs State Management

A million-token context window is large, but a persistent agent can eventually fill it. Tool results, conversation history, source material, planning notes, generated code, and intermediate research all accumulate.

Compaction addresses a different problem from raw context size. Instead of merely raising the context ceiling, it summarizes older state so the agent can continue while preserving the important information needed for later turns.

This makes Opus 4.6 historically important for long-running-agent architecture: large native context and explicit context-management infrastructure started to converge in the same generation.

For production evaluation, measure whether compaction preserves task-critical details. A session that technically continues forever is not useful if earlier constraints, decisions, or evidence disappear from the summarized state.

Prompt caching and compaction solve different economic problems. Caching makes repeated stable context cheaper; compaction reduces how much old context must remain active.

Long-session state management
  1. 01

    Keep important state

    Preserve objectives, decisions, constraints, and key evidence while older conversational detail is summarized.

  2. 02

    Control context growth

    Prevent persistent agents from indefinitely accumulating raw tool results and conversation history.

  3. 03

    Combine with caching

    Cache stable prefixes for cost efficiency while compacting older dynamic history when the session grows.

  4. 04

    Evaluate memory loss

    Replay long tasks and verify that compacted sessions still honor earlier requirements and factual dependencies.

05 / Pricing

Claude Opus 4.6 Pricing

Claude Opus 4.6 costs $5 per million input tokens and $25 per million output tokens, with separate prompt-caching rates and a 50% Batch API discount.

Agentic Cost Depends on More Than the Headline Rate

Anthropic lists 5-minute cache writes at $6.25 per million tokens, 1-hour cache writes at $10 per million, and cache reads at $0.50 per million.

These rates matter for long-running workflows because the same large context may be reused repeatedly. System instructions, repository snapshots, reference documents, tool definitions, or policy material can often remain stable across multiple turns.

Batch API requests receive a 50% discount on standard input and output pricing. This makes batch execution useful for large offline evaluations, bulk code or document processing, and workloads that do not need interactive latency.

Current Claude Opus 5.5 is cheaper at $4 input and $20 output per million tokens. That means migration from Opus 4.6 can lower the base token rate even before quality improvements are considered.

  • Input: $5.00 per 1M tokens.
  • Output: $25.00 per 1M tokens.
  • 5-minute cache write: $6.25 per 1M tokens.
  • 1-hour cache write: $10.00 per 1M tokens.
  • Cache read: $0.50 per 1M tokens.
  • Batch API: 50% discount on input and output.
Token and cache pricing

1M tokens · USD

Input
$5.00
5m cache write
$6.25
1h cache write
$10.00
Cache read
$0.50
Output
$25.00

Example: 20K input + 4K output

Input cost
$0.1000
Output cost
$0.1000
Estimated uncached total
$0.2000

06 / Context

A 1M Context Window at Standard Pricing

Claude Opus 4.6 supports a 1,000,000-token context window and up to 128,000 standard output tokens, with the full 1M window now available without a special context beta header.

Opus 4.6 Helped Make Million-Token Context the Normal Opus Baseline

When Opus 4.6 launched, its 1M context window initially arrived in beta. Anthropic later made the full 1M window generally available at standard pricing, with requests above 200K tokens working automatically.

That change is relevant for older integrations because code may still contain a legacy context beta header that is no longer necessary.

Anthropic also raised the media capacity associated with 1M context to as many as 600 images or PDF pages per request, enabling much larger multimodal document workloads.

For standard requests, maximum output is 128,000 tokens. Through the Message Batches API, Anthropic documents a beta path for up to 300,000 output tokens using the appropriate beta feature.

Large capacity does not eliminate the need for context discipline. Retrieval, caching, compaction, and relevance filtering still determine whether the million-token window is being used efficiently.

Context capacity

Context window

1,000,000

Standard max output

128,000

Input + working stateOutput limit

The full 1M context window is generally available at standard pricing. Message Batches can expose up to 300K output tokens in beta.

07 / API behavior

API Behaviors That Make Opus 4.6 Distinct

Opus 4.6 sits between older Claude APIs and the newer Opus contract, so several integration behaviors are unique enough to require explicit migration testing.

No Assistant Prefill, Legacy Tool Semantics, and Historical Fast Mode

First, Opus 4.6 does not support prefilling assistant messages. Applications that previously started an assistant response with predefined text need to use a different output-control pattern.

Second, Opus 4.6 still accepts forced tool choice and the older computer_20251124 computer-use tool. These behaviors change in newer Opus generations.

Third, Anthropic launched a research-preview fast mode for Opus 4.6 in February 2026, with output generation up to 2.5× faster at premium pricing. That fast mode was removed on June 29, 2026.

Today, sending speed: "fast" to Opus 4.6 does not provide fast inference or premium billing. The request runs at standard speed and standard rates, and usage.speed reports the actual speed used.

Finally, the 4.6 generation introduced dateless model IDs. claude-opus-4-6 identifies a pinned model version without the snapshot date format used by earlier Claude releases.

On Amazon Bedrock, Opus 4.6 is also a boundary case: it is the last model whose Bedrock identifier keeps the -v1 suffix.

Legacy integration details
  1. 01

    No assistant prefill

    Opus 4.6 does not support prefilling assistant messages, so legacy completion-shaping patterns need another implementation.

  2. 02

    Older tool contract

    Forced tool choice and the computer_20251124 interface remain valid on Opus 4.6 but change on newer Opus generations.

  3. 03

    Fast mode removed

    speed fast no longer accelerates Opus 4.6; requests run at standard speed and standard pricing.

  4. 04

    Dateless model ID

    The 4.6 generation introduced the claude-name-major-minor naming style without a dated snapshot suffix.

08 / Migration

Migrating From Claude Opus 4.6 to Opus 5.5

Moving from Opus 4.6 to Opus 5.5 changes the thinking contract, generation controls, tokenization, image processing, tool orchestration, and price, so the upgrade should be treated as an integration migration rather than a model-name replacement.

Opus 4.6 Carries More Legacy API Surface Than 4.7 or 4.8

The largest change is reasoning. Opus 4.6 can use the deprecated thinking: enabled plus budget_tokens format, can use adaptive thinking, or can run without thinking. Opus 5.5 uses adaptive thinking on every request and rejects the old manual budget format.

Anthropic’s migration guide also instructs Opus 4.6 users to remove temperature, top_p, and top_k when moving to Opus 5.5.

Newer Opus generations use updated tokenization, so client-side token estimates, max_tokens, and compaction thresholds should be recalibrated.

High-resolution image processing is another cumulative change introduced after 4.6. Image-heavy applications should re-budget token usage and downsample images when full visual fidelity is unnecessary.

Tool orchestration also changes across the generations: forced tool choice, computer-use versions, mid-conversation tool changes, system messages, and thinking-block handling all need explicit testing.

Finally, Opus 5.5 lowers standard pricing from $5/$25 to $4/$20 and defaults to medium effort instead of high.

Opus 4.6 → current Opus
  1. 01

    Replace legacy thinking budgets

    Remove thinking enabled plus budget_tokens and migrate to adaptive thinking with effort; current Opus no longer accepts the manual format.

  2. 02

    Remove incompatible sampling controls

    Review temperature, top_p, and top_k because current Opus migration guidance requires removing them.

  3. 03

    Recount tokens and images

    Re-test max_tokens, compaction thresholds, client-side token estimates, and image budgets under newer tokenization and vision behavior.

  4. 04

    Replay tool and conversation flows

    Test forced tools, computer use, thinking blocks, cached histories, dynamic tools, system messages, and long-running sessions before production cutover.

09 / Capabilities

Claude Opus 4.6 Capabilities

Claude Opus 4.6 combines multimodal input, 1M context, adaptive and legacy thinking modes, tool use, prompt caching, batch processing, and long-session context management in a legacy frontier model.

A Feature-Rich Compatibility Baseline

The model accepts text and images and produces text. Adaptive thinking is the recommended reasoning path, but legacy extended thinking remains available for older integrations.

Tool use supports complex agentic workflows, and prompt caching reduces repeated processing cost for stable context. The model also supports large Batch API outputs and was the launch model for Anthropic’s compaction API.

Its integration surface is more permissive than current Opus in several ways: thinking can be absent or explicitly disabled, forced tool choice is available, and the older computer-use tool is accepted.

At the same time, some newer features are absent or behave differently. Assistant prefilling is unsupported, high-resolution vision arrived in Opus 4.7, and later models introduce additional dynamic orchestration controls.

Supported capabilities
  • Text

    Accept text input and generate text output.

    Supported
  • Image input

    Analyze supported images together with text and long-context material.

    Supported
  • Adaptive thinking

    Recommended reasoning path controlled with effort.

    Supported
  • Legacy extended thinking

    Manual budget_tokens mode still works but is deprecated and removed from Opus 4.7 onward.

    Supported
  • Tool use

    Supports agentic tool workflows under the older forced-tool and computer-use contract.

    Supported
  • Prompt caching

    Reduce repeated processing cost for stable large-context prefixes.

    Supported
  • Batch API

    Batch processing receives a 50% discount and can expose up to 300K output tokens in beta.

    Supported

10 / Evaluation

Claude Opus 4.6 Strengths and Limitations

Claude Opus 4.6 remains valuable as a compatibility and historical baseline because it combines modern-scale context with both old and new reasoning APIs, but its model quality, pricing, and orchestration contract have been superseded by newer Opus generations.

Strengths

  • Dual thinking compatibility

    Supports both adaptive thinking and the deprecated manual extended-thinking format, which is useful for legacy integrations.

  • 1M context and 128K output

    Large working capacity supports repositories, documents, persistent conversations, tool history, and substantial generated artifacts.

  • Long-horizon agent baseline

    Anthropic launched Opus 4.6 specifically for complex agentic tasks and long-running work.

  • Compaction-era architecture

    The model launched alongside server-side compaction, making it a useful reference for persistent-agent context management.

What to consider

  • Legacy model

    Anthropic recommends migrating to Claude Opus 5.5 for improved current-generation performance.

  • Higher current Opus pricing

    $5/$25 standard pricing is higher than Opus 5.5 at $4/$20.

  • Legacy reasoning API

    Manual budget_tokens still works but is deprecated, and migration to current Opus requires removing it.

  • Historical fast mode is gone

    speed fast no longer accelerates Opus 4.6; requests now run at standard speed and standard billing.

Understand the transition before migrating

Compare Claude Opus 4.6 with newer Opus models

Replay long-horizon agents, reasoning-heavy prompts, tool calls, cached contexts, compaction workflows, and legacy thinking configurations to see how current Opus models change quality, cost, latency, and integration behavior.

Start Free

Anthropic keeps Claude Opus 4.6 available as a legacy model but recommends migrating to Claude Opus 5.5 for improved performance.

Common Questions

What is Claude Opus 4.6?

Claude Opus 4.6 is an Anthropic legacy frontier model released on February 5, 2026 for complex agentic tasks and long-horizon work. It is notable for supporting both adaptive thinking and the deprecated manual extended-thinking API.

How much does Claude Opus 4.6 cost?

Claude Opus 4.6 costs $5 per 1M input tokens and $25 per 1M output tokens. Five-minute cache writes cost $6.25/MTok, one-hour cache writes cost $10/MTok, and cache reads cost $0.50/MTok. Batch input and output receive a 50% discount.

What is the Claude Opus 4.6 context window?

Claude Opus 4.6 has a 1,000,000-token context window. The full 1M context is generally available at standard pricing without a context beta header.

What is the maximum output of Claude Opus 4.6?

Standard requests support up to 128,000 output tokens. Anthropic also documents a beta path for up to 300,000 output tokens through the Message Batches API.

What is Claude Opus 4.6’s knowledge cutoff?

Anthropic lists May 2025 as the reliable knowledge cutoff and August 2025 as the training-data cutoff.

Does Claude Opus 4.6 support adaptive thinking?

Yes. Anthropic recommends adaptive thinking with effort controls. Opus 4.6 defaults to high effort when effort is used, and thinking itself is off when the request does not ask for it.

Does Claude Opus 4.6 still support budget_tokens?

Yes. Manual extended thinking with thinking type enabled and budget_tokens still works on Opus 4.6, but Anthropic marks it deprecated. Starting with Opus 4.7, that format is rejected.

What effort levels does Claude Opus 4.6 support?

Claude Opus 4.6 supports effort controls including low, medium, high, and max, with high as the API default. Unlike later Opus models, xhigh is not listed as available for Opus 4.6.

Does Claude Opus 4.6 support fast mode?

Not anymore. Anthropic launched fast mode for Opus 4.6 in February 2026 but removed it on June 29, 2026. Requests that still send speed fast now run at standard speed and standard pricing.

Does Claude Opus 4.6 support assistant prefilling?

No. Anthropic explicitly documents that Opus 4.6 does not support prefilling assistant messages.

What is compaction in Claude Opus 4.6 workflows?

Anthropic launched its compaction API alongside Opus 4.6. Compaction summarizes older server-side context so long-running conversations and agents can continue without retaining every previous token verbatim.

How is Claude Opus 4.6 different from Claude Opus 4.7?

Opus 4.6 supports both adaptive thinking and deprecated manual budget_tokens thinking. Opus 4.7 removes the manual thinking path, introduces a new tokenizer, adds higher-resolution vision, and changes several agent behaviors while keeping the same $5/$25 base pricing.

How is Claude Opus 4.6 different from Claude Opus 5.5?

Opus 4.6 can run without thinking and still accepts the legacy manual thinking API, forced tool choice, and older computer-use behavior. Opus 5.5 uses always-on adaptive thinking, changes tool and conversation semantics, and lowers standard pricing to $4 input and $20 output per million tokens.

Is Claude Opus 4.6 deprecated?

Anthropic lists Claude Opus 4.6 as Active (legacy). It was released February 5, 2026, with retirement not sooner than February 5, 2027, and Anthropic recommends migrating to Claude Opus 5.5.

Should I use Claude Opus 4.6 for a new application?

For new development, evaluate the current Opus generation first. Claude Opus 4.6 is most useful for existing integrations, regression testing, legacy thinking compatibility, and migration baselines.

Model information

Last updated

Specifications, pricing, lifecycle, thinking behavior, effort controls, compaction history, context limits, fast-mode status, and migration guidance on this page are based on Anthropic’s official Claude Opus 4.6 documentation, thinking documentation, migration guide, model-versioning guide, and Claude Platform release notes.

Claude Opus 4.6 — Pricing, 1M Context, Extended vs Adaptive Thinking | EidoStack