Anthropic legacy frontier model
Claude Opus 4.8
A late Claude 4.x frontier model for long-horizon agentic work, knowledge work, vision, and memory-heavy workflows, with a 1M-token context window, 128K output, optional adaptive thinking, and mature tool-oriented APIs.
- Context window
- 1M
- tokens
- Max output
- 128K
- tokens
- Input
- $5.00
- per 1M tokens
- Cache read
- $0.50
- per 1M tokens
- Output
- $25.00
- per 1M tokens
01 / Overview
What Claude Opus 4.8 Is
Claude Opus 4.8 is a late Claude 4.x frontier model designed for long-horizon agentic work, complex knowledge tasks, vision, and memory-heavy workflows before the Claude 5 generation became Anthropic’s primary frontier line.
A Legacy Model With a Modern-Scale Context Window
Anthropic released Claude Opus 4.8 on May 28, 2026 as its most capable model at the time. It brought a 1,000,000-token context window by default across the Claude API and major cloud platforms, plus a 128,000-token standard output ceiling.
Those capacities made Opus 4.8 suitable for large repositories, lengthy research material, long conversations, tool histories, and other workflows that accumulate substantial working state.
Today, Anthropic lists the model as Active (legacy). It remains available, but the provider recommends moving to Claude Opus 5.5 for improved performance.
That makes Opus 4.8 especially useful as a compatibility and evaluation baseline for applications that were built before Claude 5 changed the thinking and orchestration contract.
- Model ID:
claude-opus-4-8. - Released May 28, 2026.
- Active legacy status.
- Retirement not sooner than May 28, 2027.
- Text and image input with text output.
- 1M context and 128K standard output.
- Provider
- Anthropic
- Family
- Claude Opus 4.8
- Model ID
- claude-opus-4-8
- Knowledge cutoff
- Jan 2026
- Input modalities
- Text, Image
- Output modality
- Text
- Lifecycle
- Active · legacy
02 / Agentic work
Long-Horizon Agents, Coding, and Knowledge Work
Anthropic’s model-specific prompting guidance highlights long-horizon agentic work, knowledge work, vision, and memory tasks as particular strengths of Claude Opus 4.8.
Evaluate the Trajectory, Not the First Turn
A long-running agent can spend dozens of steps inspecting files, calling tools, refining a plan, resolving failures, and updating its internal state. For these workloads, a single-response benchmark is incomplete.
The useful question is whether the model can keep the task on course. Measure completion rate, unnecessary tool calls, repeated work, recovery after errors, human interventions, and whether the final output satisfies the original goal.
Opus 4.8 is also tuned to calibrate response length based on task complexity. Simple lookups can produce concise answers, while open-ended analysis may become substantially more detailed.
That behavior matters when migrating. A downstream interface, parser, or reviewer may have assumptions about how much explanation the model returns. Re-evaluate response length as part of the migration rather than treating verbosity as cosmetic.
- 01
Repository-scale coding
Inspect many files, reason across dependencies, call tools, modify code, and continue until tests or validation succeed.
- 02
Knowledge work
Synthesize long documents, technical references, policies, or research while preserving the task objective.
- 03
Vision-heavy workflows
Combine screenshots, document images, diagrams, and textual instructions inside one larger reasoning process.
- 04
Long-running agents
Maintain state across extended sequences of planning, tool use, correction, and verification.
03 / Memory & context
1M Context for Memory-Heavy Workflows
Claude Opus 4.8 provides a 1,000,000-token context window and supports up to 128,000 output tokens, giving agentic systems room for both large source material and accumulated working state.
Context Is Working Memory, Not Just Document Size
For an agent, context can contain system instructions, source code, previous messages, tool results, screenshots, research evidence, plans, memory entries, and intermediate output.
A million-token ceiling reduces the need to discard state aggressively, but it does not eliminate context-management decisions. Irrelevant history still consumes tokens and can make important instructions harder to distinguish.
Prompt caching is therefore particularly valuable with Opus 4.8. Stable prefixes such as system prompts, large repositories, policies, or reference documents can be cached and reused across requests instead of being billed repeatedly at the full standard input rate.
Anthropic also offers a beta path for up to 300,000 output tokens through the Message Batches API. Standard requests retain the 128K output ceiling.
Context window
1,000,000
Standard max output
128,000
The Batch API can support up to 300K output tokens in beta; standard requests use the 128K output limit.
04 / Pricing
Claude Opus 4.8 Pricing
Claude Opus 4.8 costs $5 per million input tokens and $25 per million output tokens, with separate prompt-caching rates and a 50% Batch API discount.
Legacy Pricing Is Now Part of the Migration Case
Anthropic lists 5-minute cache writes at $6.25 per million tokens, 1-hour cache writes at $10 per million, and cache reads at $0.50 per million.
Those rates can materially reduce repeated-context cost for long-running agents. A repository, policy set, or large system prompt may be expensive to ingest once but relatively cheap to reuse through cache reads.
However, Claude Opus 5.5 is cheaper at standard token rates: $4 input and $20 output per million tokens. That means an Opus 4.8 migration can improve economics even before considering quality or latency changes.
For agentic workloads, compare cost per completed task. Include tool turns, retries, cache writes, cache reads, output, failures, and any fallback calls.
- Input: $5.00 per 1M tokens.
- Output: $25.00 per 1M tokens.
- 5-minute cache write: $6.25 per 1M tokens.
- 1-hour cache write: $10.00 per 1M tokens.
- Cache read: $0.50 per 1M tokens.
- Batch API: 50% discount on input and output.
1M tokens · USD
- Input
- $5.00
- 5m cache write
- $6.25
- 1h cache write
- $10.00
- Cache read
- $0.50
- Output
- $25.00
Example: 20K input + 4K output
- Input cost
- $0.1000
- Output cost
- $0.1000
- Estimated uncached total
- $0.2000
05 / Thinking
Adaptive Thinking Is Available, but Not Always On
Claude Opus 4.8 supports adaptive thinking with high as the default effort level when thinking is used, but requests that omit the thinking field run without thinking.
This Is a Major Difference From Claude 5
Opus 4.8 sits at an important API transition point.
The model supports adaptive thinking and effort controls, but thinking is not automatically activated for every request. Anthropic’s migration guide states that an Opus 4.8 request with no thinking field runs without thinking.
It also accepts thinking: {"type": "disabled"}. That flexibility disappears in Claude Opus 5.5, where thinking is always present and cannot be disabled.
This difference can affect both quality and economics. A production integration that never enabled thinking on Opus 4.8 may see a significant behavior change after a direct model-ID swap to Opus 5.5.
Before migrating, identify which workloads intentionally used thinking, which ran without it, and how much output budget was available. On newer models, thinking and visible response text share the output limit.
Request without thinking
On Opus 4.8, a request that omits the thinking field runs without thinking.
Reasoning-heavy request
Enable adaptive thinking and calibrate effort when difficult planning, coding, or analysis benefits from additional reasoning.
06 / System messages
Mid-Conversation System Messages
Claude Opus 4.8 introduced the ability to add system-role messages after a user turn, allowing applications to update instructions during a long-running session without rebuilding the full cached conversation prefix.
Change Instructions Without Throwing Away Earlier Cache Hits
Long-running agents frequently need new instructions as the task evolves. A supervisor may narrow scope, update a policy, change priorities, or tell the model how to handle a newly discovered condition.
Before mid-conversation system messages, an application might need to reconstruct the system prompt or rebuild history to introduce those changes. That can invalidate prompt-cache prefixes and complicate state management.
On Opus 4.8, supported API paths can insert role: "system" messages after a user turn, subject to Anthropic’s placement rules. Anthropic launched the capability without requiring a beta header.
This is a practical model-specific feature worth testing in persistent agents because it affects both orchestration design and cache efficiency.
- 01
Preserve prior context
Keep the existing conversation instead of rebuilding the entire history when instructions change.
- 02
Protect cache efficiency
Update later-turn guidance while preserving eligible prompt-cache hits on earlier content.
- 03
Supervise long-running agents
Adjust scope, policy, priorities, or task constraints as new information appears.
- 04
Test placement rules
Validate the exact message sequence your integration uses because system-role insertion follows provider-defined placement constraints.
07 / Migration
Migrating From Claude Opus 4.8 to Opus 5.5
Anthropic recommends migrating from Opus 4.8 to Opus 5.5, but several API behaviors change enough that the upgrade should be treated as a compatibility project rather than a pure model-name replacement.
Thinking, Tools, Caching, and Capacity Need Retesting
The most important change is thinking. Opus 4.8 can run with thinking off; Opus 5.5 always uses thinking. Requests that explicitly disable thinking or use older manual thinking-budget modes are incompatible with Opus 5.5.
Forced tool choice also changes. Opus 4.8 accepts tool-choice modes that force any tool or a named tool, while Opus 5.5 requires a different orchestration pattern.
The prompt-cache minimum becomes smaller on Opus 5.5: 512 tokens instead of 1,024 on Opus 4.8. This can make shorter stable prompts cacheable after migration.
Priority Tier is another difference. Anthropic’s migration documentation states that Opus 5.5 does not support Priority Tier, while Opus 4.8 does.
Finally, re-baseline cost and latency. Opus 5.5 lowers standard pricing from $5/$25 to $4/$20, defaults to medium effort, and introduces different thinking behavior that can change total output consumption.
- 01
Audit thinking settings
Find requests that omit thinking or disable it; Opus 5.5 runs with thinking on and rejects disabled or manual-budget thinking modes.
- 02
Update tool orchestration
Replace forced any-tool or named-tool patterns where necessary and retest agent loops.
- 03
Revisit caching and capacity
Opus 5.5 lowers the cacheable-prompt minimum from 1,024 to 512 tokens and does not support Priority Tier.
- 04
Re-baseline production metrics
Measure completion quality, latency, output usage, cache behavior, refusal handling, and cost under the new model.
08 / Capabilities
Claude Opus 4.8 Capabilities
Opus 4.8 combines multimodal input, large context, adaptive thinking, tool use, prompt caching, long output, and mature long-running-agent behavior in a legacy Claude 4.x model.
A Mature Pre-Claude-5 Integration Surface
The model accepts text and images and returns text. Its prompting guidance emphasizes long-horizon agentic work, knowledge work, vision, and memory tasks.
Tool use supports agent orchestration, and Opus 4.8 remains compatible with forced tool-choice patterns that later Opus 5.5 changes. Prompt caching helps reduce repeated large-context cost.
Anthropic also publicly documents stop_details for refusal responses introduced alongside Opus 4.8. The field can include a refusal category such as cyber, bio, or null, plus a human-readable explanation, which allows applications to route different refusal classes differently.
That makes refusal handling an application concern rather than a generic error state. Production systems should treat stop_reason: "refusal" as a distinct outcome and define the appropriate user-facing or fallback behavior.
- Supported
Text
Accept text input and produce text output.
- Supported
Image input
Analyze screenshots, documents, diagrams, and other supported visual material.
- Supported
Adaptive thinking
Available with effort controls, but unlike Claude 5 it is not always on.
- Supported
Tool use
Supports agentic tool workflows, including forced tool-choice patterns available to this generation.
- Supported
Prompt caching
Reduce repeated cost for large stable prompts and long-session context.
- Supported
Batch API
Batch processing receives a 50% input/output discount and can expose up to 300K output in beta.
- Supported
Refusal metadata
stop_details can provide refusal category and explanation for application-level routing.
09 / Evaluation
Claude Opus 4.8 Strengths and Limitations
Claude Opus 4.8 remains a capable legacy baseline for long-context agents, memory, knowledge work, and vision, but Claude 5 offers a newer reasoning contract, lower current Opus pricing, and improved frontier performance.
Strengths
1M context and 128K output
Large working-state capacity supports repositories, research material, persistent conversations, and substantial generated artifacts.
Long-horizon agentic behavior
Anthropic specifically highlights agentic work, knowledge work, vision, and memory as Opus 4.8 strengths.
Flexible thinking contract
Applications can run without thinking or enable adaptive reasoning selectively, unlike later always-on Claude 5 models.
Dynamic system instructions
Mid-conversation system messages can update long-running sessions while preserving eligible earlier cache hits.
What to consider
Legacy status
Anthropic keeps Opus 4.8 active but recommends migrating to Claude Opus 5.5 for improved performance.
Higher price than current Opus
$5/$25 standard pricing is higher than Opus 5.5 at $4/$20.
Migration changes reasoning semantics
Opus 5.5 always uses thinking, so integrations that intentionally ran Opus 4.8 without thinking need explicit re-evaluation.
Older orchestration contract
Forced tool choice, Priority Tier, caching minimums, and some computer-use assumptions differ from the current Opus generation.
Preserve a late Claude 4.x baseline
Compare Claude Opus 4.8 with current Opus models
Replay your agentic coding, research, vision, memory, long-context, tool-use, and refusal-sensitive workloads to measure whether Opus 5.5 improves completion quality, latency, cost, and operational simplicity.
Start FreeClaude Opus 4.8 is still active, but Anthropic classifies it as legacy and recommends migrating to Claude Opus 5.5.
Common Questions
What is Claude Opus 4.8?
Claude Opus 4.8 is an Anthropic legacy frontier model released on May 28, 2026 for long-horizon agentic work, knowledge work, vision, memory, and other complex workflows.
How much does Claude Opus 4.8 cost?
Claude Opus 4.8 costs $5 per 1M input tokens and $25 per 1M output tokens. Five-minute cache writes cost $6.25/MTok, one-hour cache writes cost $10/MTok, and cache reads cost $0.50/MTok. Batch input and output receive a 50% discount.
What is the Claude Opus 4.8 context window?
Claude Opus 4.8 has a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.
Can Claude Opus 4.8 generate more than 128K tokens?
Anthropic documents a beta maximum of up to 300,000 output tokens through the Message Batches API. Standard requests use the 128K output limit.
What is Claude Opus 4.8’s knowledge cutoff?
Anthropic lists January 2026 as both the reliable knowledge cutoff and training-data cutoff.
Does Claude Opus 4.8 support images?
Yes. Claude Opus 4.8 accepts text and image input and generates text output.
Does Claude Opus 4.8 use adaptive thinking?
Yes, but thinking is not always on. Anthropic’s migration documentation states that Opus 4.8 requests that omit the thinking field run without thinking. The model supports adaptive thinking with high as its default effort level when reasoning is used.
How is Claude Opus 4.8 different from Claude Opus 5.5?
Opus 4.8 can run without thinking, accepts older forced tool-choice patterns, has a 1,024-token prompt-cache minimum, supports Priority Tier, and costs $5/$25. Opus 5.5 uses always-on thinking, lowers standard pricing to $4/$20, lowers the cacheable-prompt minimum to 512 tokens, and changes several orchestration behaviors.
What are mid-conversation system messages in Claude Opus 4.8?
They let supported integrations add system-role instructions after a user turn during an existing conversation. This can update agent guidance without rebuilding the entire conversation and can preserve eligible prompt-cache hits on earlier content.
What is stop_details in Claude Opus 4.8?
Anthropic documents stop_details on refusal responses as metadata that can include a refusal category such as cyber, bio, or null together with a human-readable explanation. Applications can use it to route refusal outcomes appropriately.
Is Claude Opus 4.8 deprecated?
Anthropic lists Claude Opus 4.8 as Active (legacy), not retired. It was released May 28, 2026, with retirement not sooner than May 28, 2027, but Anthropic recommends migrating to Claude Opus 5.5.
Should I use Claude Opus 4.8 for a new application?
For a new evaluation, Anthropic recommends the current Opus 5.5 model. Opus 4.8 is most useful for existing integrations, compatibility testing, historical evaluation, or preserving a production baseline before migration.
Model information
Last updated
Specifications, pricing, thinking behavior, prompting characteristics, lifecycle, migration differences, and API features on this page are based on Anthropic’s official Claude Opus 4.8 overview, prompting guide, migration guide, and release notes.