Anthropic model

Claude Sonnet 5.5

Anthropic’s fast frontier model for production coding, tool-heavy agents, and professional knowledge work, combining a 1M-token context window with lower pricing and latency than the Opus tier.

Context window
1M
tokens
Max output
128K
tokens
Input
$2.00
per 1M tokens
Cache read
$0.20
per 1M tokens
Output
$10.00
per 1M tokens

01 / Overview

What Claude Sonnet 5.5 Is

Claude Sonnet 5.5 is Anthropic’s current fast frontier model, designed to combine strong intelligence with lower latency and lower token cost than the Opus and Fable tiers.

A Frontier Model Designed to Be Used More Often

Anthropic describes Sonnet 5.5 as its best combination of speed and intelligence.

That positioning matters in production. The highest-capability model is not automatically the best default for every request. A SaaS product may make thousands or millions of model calls, and a modest reduction in per-request cost or latency can compound across the entire workload.

Sonnet 5.5 keeps the same large-scale technical envelope as the current Opus tier: a 1,000,000-token context window, up to 128,000 standard output tokens, text and image input, adaptive thinking, prompt caching, and advanced tool workflows.

The difference is economics and latency. Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, while Anthropic classifies its comparative latency as fast.

Anthropic released Claude Sonnet 5.5 on September 28, 2026 and lists it as Active (latest), with retirement not scheduled sooner than September 28, 2027.

  • Model ID: claude-sonnet-5-5.
  • 1M-token context window.
  • 128K standard maximum output.
  • Text and image input with text output.
  • Adaptive thinking enabled by default.
  • Fast comparative latency.
  • $2 input / $10 output per million tokens.
Model profile
Provider
Anthropic
Family
Claude Sonnet 5.5
Model ID
claude-sonnet-5-5
Knowledge cutoff
Jun 2026
Input modalities
Text, Image
Output modality
Text
Comparative latency
Fast

02 / Production fit

When Sonnet 5.5 Makes Sense as the Production Default

Sonnet 5.5 is most compelling when an application needs frontier-level behavior frequently enough that Opus pricing or latency would materially affect unit economics.

Optimize for the Cheapest Model That Reliably Passes the Task

A useful production routing strategy does not ask which model is theoretically strongest. It asks which model clears the required quality threshold for a specific workload.

Sonnet 5.5 can be a strong default for coding assistance, document analysis, customer-facing professional workflows, research synthesis, structured tool use, and multi-step agents where both quality and responsiveness matter.

More difficult tasks can then escalate to Opus or Fable only when evaluation shows that Sonnet fails too often.

This routing pattern can lower total inference cost because the expensive tier is reserved for the minority of requests that actually need it.

The correct decision should come from workload-specific evals. Measure accepted-result rate, retries, human intervention, latency, output length, and end-to-end task cost rather than comparing benchmark scores in isolation.

Production routing
  1. 01

    Default path

    Route routine but meaningful coding, research, analysis, and tool workflows to Sonnet when it consistently meets the acceptance threshold.

  2. 02

    Escalation path

    Move unusually difficult or long-horizon failures to Opus or Fable rather than paying premium-model prices on every request.

  3. 03

    Latency-sensitive path

    Use lower effort where fast response time matters and evaluation shows that additional reasoning does not improve the outcome.

  4. 04

    Measure the whole task

    Compare cost per accepted result, including retries, tool turns, cache activity, output length, and fallbacks.

03 / Agentic coding

Agentic Coding Without Opus-Level Token Pricing

Sonnet 5.5 is particularly relevant for coding agents that need strong repository reasoning and repeated tool use but must remain economical enough for frequent production execution.

Evaluate Completion Efficiency, Not Only Code Quality

An agentic coding workload includes far more than generating a function.

The model may inspect repository structure, read files, search symbols, plan a change, call tools, edit several files, run tests, investigate failures, and repeat until the implementation is complete.

For these tasks, the useful comparison is trajectory efficiency:

  • How often does the agent finish successfully?
  • How many tool calls does it need?
  • How much context does it accumulate?
  • How often does it repeat work?
  • How many tokens are consumed before the tests pass?
  • When does a human need to intervene?

Anthropic’s guidance for Sonnet 5.5 recommends starting well-specified agentic coding and multistep tool workloads around medium effort and moving toward high for harder or longer tasks.

That is economically important. A lower-cost model running a shorter, cleaner trajectory can outperform a stronger model on total cost even when the stronger model has a higher one-shot success rate.

Coding and tool workflows
  1. 01

    Feature implementation

    Navigate a repository, modify multiple files, call tools, run tests, and keep iterating until the requested feature is complete.

  2. 02

    Debugging

    Trace failures through code, logs, configuration, and tool output while keeping the original problem in working context.

  3. 03

    Code review

    Inspect implementation details and return actionable findings without making every review request an Opus-tier call.

  4. 04

    High-volume engineering automation

    Run repeated coding and maintenance jobs where per-task latency and inference cost compound quickly.

04 / Pricing

Claude Sonnet 5.5 Pricing

Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, with the same base pricing as Claude Sonnet 5.

The Economics Are Designed for Frequent Production Use

Prompt caching can reduce repeated-context cost further.

Anthropic lists 5-minute cache writes at $2.50 per million tokens, 1-hour cache writes at $4 per million, and cache reads at $0.20 per million.

The model also supports Batch API processing with a 50% discount on standard input and output. That can be useful for offline evaluations, bulk document processing, code analysis, or other workloads that do not need interactive latency.

Sonnet 5.5 is half the standard token price of Opus 5.5: $2/$10 versus $4/$20. Fable 5.1 is substantially more expensive at $10/$50.

This is why model routing matters. If Sonnet reaches the same application acceptance threshold as a premium tier on most requests, the cost difference can be significant at scale.

  • Input: $2.00 per 1M tokens.
  • Output: $10.00 per 1M tokens.
  • 5-minute cache write: $2.50 per 1M tokens.
  • 1-hour cache write: $4.00 per 1M tokens.
  • Cache read: $0.20 per 1M tokens.
  • Batch API: 50% discount on input and output.
Token and cache pricing

1M tokens · USD

Input
$2.00
5m cache write
$2.50
1h cache write
$4.00
Cache read
$0.20
Output
$10.00

Example: 20K input + 4K output

Input cost
$0.0400
Output cost
$0.0400
Estimated uncached total
$0.0800

05 / Context

1M Context and 128K Standard Output

Claude Sonnet 5.5 provides a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.

Large Enough for Repository, Research, and Agent State

A million-token context can hold a large repository slice, long documents, screenshots, conversation history, tool results, system instructions, retrieved evidence, and intermediate agent state.

This makes Sonnet 5.5 suitable for workloads where the application wants a large working set without immediately escalating to a more expensive model tier.

Anthropic also documents up to 300,000 output tokens through the Message Batches API in beta. That higher limit is specific to batch processing; standard requests retain the 128K output ceiling.

Prompt caching becomes especially relevant with large contexts. Sonnet 5.5 lowers the minimum cacheable prompt to 512 tokens, compared with 1,024 tokens on Sonnet 5.

A larger context window still does not remove the need for retrieval and state management. Sending irrelevant history consumes money and can reduce signal quality even when it technically fits.

Context capacity

Context window

1,000,000

Standard max output

128,000

Input + working stateOutput limit

Message Batches can expose up to 300K output tokens in beta. Prompt caching supports a 512-token minimum cacheable prompt.

06 / Thinking

Adaptive Thinking With a `between_tools` Low-Compute Path

Claude Sonnet 5.5 uses adaptive thinking by default, but unlike the always-on Opus 5.5 contract, it also provides between_tools as a lower-compute mode that removes up-front thinking while preserving short reasoning updates during tool workflows.

Thinking Can Be Tuned for Latency-Sensitive Work

On the Claude API, the default effort level is high.

For well-specified agentic coding and multistep tool use, Anthropic recommends starting around medium and moving to high for more difficult or longer tasks. For chat and latency-sensitive workloads, medium or low can be a better starting point.

Sonnet 5.5 also introduces a specific alternative to disabled thinking: thinking: {"type": "between_tools"}.

With between_tools, the model does not perform up-front thinking before the first response, but can still generate progress-oriented thinking blocks between tool calls. Anthropic describes it as the lowest thinking setting on the model.

between_tools works at low, medium, and high effort. At xhigh or max, the request must use adaptive thinking instead.

The older thinking: disabled mode returns a 400 error, as does manual budget_tokens thinking.

Adaptive thinking modes
lowest thinking modelowmediumhigh · API defaultxhighmax

Latency-sensitive interaction

Start with low or medium effort, or use between_tools when the workflow uses tools but does not need up-front reasoning.

Hard agentic task

Use adaptive thinking at high or above when deeper reasoning materially improves completion quality.

07 / Tools & state

Dynamic Tools, Compaction, and Long-Running State

Sonnet 5.5 adds several controls that are particularly useful for persistent agents: mid-conversation system messages, mid-conversation tool changes, on-demand compaction, and inline tool definitions.

Modify the Agent Without Rebuilding the Conversation

Anthropic supports mid-conversation system messages on Sonnet 5.5. This lets an application update instructions during a session without editing earlier turns.

Mid-conversation tool changes are available in beta. An application can change the available tool set as a task progresses instead of exposing every tool from the beginning.

Anthropic also documents compact on demand in beta. The API can return a signed compaction block that summarizes the prior conversation so the application can replace older raw turns while preserving the working state needed to continue.

Another beta feature allows tool definitions to be introduced inside a mid-conversation system message. This can add a tool, modify its schema, or move to a newer server-tool version without rewriting the top-level tool list and losing earlier prompt-cache continuity.

These controls make Sonnet 5.5 especially relevant for agents whose state, permissions, and available actions evolve over time.

Dynamic agent controls
  • Mid-conversation system messages

    Update instructions during a session without editing the earlier conversation history.

    Supported
  • Mid-conversation tool changes

    Beta support lets applications add or remove tools as the agent progresses.

    Supported
  • Compaction on demand

    Beta compaction can replace older conversation history with a signed summarized state.

    Supported
  • Inline tool definitions

    Beta system-message blocks can add or change full tool definitions without rebuilding the top-level tool list.

    Supported
  • Prompt caching

    A 512-token minimum makes shorter reusable prefixes eligible for caching.

    Supported
  • Computer use

    On the Claude API and Google Cloud, Sonnet 5.5 uses the computer_toolset_20260801 interface.

    Supported

08 / Refusals

Refusals Are a Structured Application Outcome

Claude Sonnet 5.5 can decline requests with stop_reason: "refusal" and structured stop_details, which lets applications handle different safety outcomes explicitly instead of treating them as generic failures.

Production Systems Should Route Refusals Deliberately

Anthropic documents five safeguard categories for Sonnet 5.5:

  • cyber
  • bio
  • frontier_llm
  • reasoning_extraction
  • general_harms

A refusal still returns HTTP 200. The application should inspect the stop reason and metadata rather than assume that a successful HTTP response always contains an ordinary completion.

Anthropic also provides beta server-side fallback behavior for some refusal classes. For example, configured fallback can retry certain cyber or frontier_llm declines on Claude Sonnet 5.

This means refusal handling belongs in production orchestration. Track refusal rate by workload, distinguish policy outcomes from ordinary model failures, and decide whether a safe fallback, user-facing explanation, or no retry is appropriate.

Refusal-aware routing
  1. 01

    Inspect stop_reason

    Treat refusal as a distinct successful API response shape rather than a transport failure.

  2. 02

    Read stop_details

    Use the documented safeguard category to understand what class of request was declined.

  3. 03

    Configure safe fallback

    Use provider-supported or application-level fallbacks only for categories where retry behavior is appropriate.

  4. 04

    Measure refusal rate

    Include refusals in production evaluation because they affect task completion, routing economics, and user experience.

09 / Sonnet 5 → 5.5

Migrating From Claude Sonnet 5 to Sonnet 5.5

Sonnet 5.5 keeps Sonnet 5 pricing and tokenization, but several thinking, tool, computer-use, and conversation-state behaviors change enough to require integration testing.

Same Economics Does Not Mean Drop-In Compatibility

The tokenizer is the same as Sonnet 5, so identical text should produce the same token counts. Standard prices also remain $2 input and $10 output per million tokens.

The breaking changes are mostly behavioral and orchestration-related.

First, thinking: disabled is replaced by between_tools when an application wants to remove up-front thinking.

Second, forced tool use is no longer supported. tool_choice: any and named-tool forcing return errors.

Third, thinking blocks are bound to the model and conversation. Editing the system prompt, tools, or earlier messages can invalidate preserved Sonnet 5.5 thinking blocks, so append-only conversation design is safer.

Fourth, computer use on the Claude API and Google Cloud moves from the older computer_20251124 tool to computer_toolset_20260801.

Fifth, some advisor-model pairings that worked with Sonnet 5 are not accepted by Sonnet 5.5.

Finally, progress text between tool calls may now arrive in thinking blocks. Applications that stream agent status to users should test thinking.display or between_tools behavior explicitly.

Migration differences
  1. 01

    Replace disabled thinking

    Use between_tools at supported effort levels instead of thinking disabled.

  2. 02

    Remove forced tool selection

    Move any-tool or named-tool forcing to automatic selection plus strict schemas or structured output patterns.

  3. 03

    Keep histories append-only

    Avoid editing earlier prompts, tools, or messages when preserved thinking blocks are replayed.

  4. 04

    Update agent integrations

    Retest computer use, advisor pairings, progress streaming, dynamic tools, and conversation-state handling.

10 / Evaluation

Claude Sonnet 5.5 Strengths and Limitations

Sonnet 5.5 is a strong production-oriented frontier model because it combines 1M context and advanced agent controls with fast comparative latency and half the base token price of Opus 5.5.

Strengths

  • Strong price-performance

    $2/$10 standard pricing makes frontier-level coding, analysis, and tool workflows substantially cheaper than current Opus or Fable tiers.

  • Fast frontier latency

    Anthropic classifies Sonnet 5.5 as fast, making it suitable for interactive and high-volume production paths.

  • 1M context and 128K output

    Large working-state capacity supports repositories, long documents, images, agent history, and substantial generated artifacts.

  • Advanced agent-state controls

    Mid-conversation instructions, dynamic tools, compaction, prompt caching, and between_tools reasoning support sophisticated long-running workflows.

What to consider

  • Not Anthropic’s maximum capability tier

    Difficult long-horizon tasks may still justify escalation to Opus 5.5 or Fable 5.1 when Sonnet fails the required evaluation threshold.

  • Forced tool use is unavailable

    Applications that depend on guaranteed named-tool invocation need a different orchestration pattern.

  • Thinking state needs careful handling

    Preserved thinking blocks are tied to model and conversation state, so editing historical context can create compatibility errors.

  • Several advanced controls are beta

    Per-message effort, dynamic tool changes, compaction on demand, and inline tool definitions may evolve and should be isolated behind stable application abstractions.

Optimize the production default

Evaluate Claude Sonnet 5.5 on your real workload

Run the same coding, tool-use, research, document, and long-context tasks across Sonnet and Opus, then compare accepted-result quality, latency, token usage, cache efficiency, and total cost per completed task.

Start Free

For many production systems, the useful question is not whether Sonnet is the strongest model available, but whether it reaches the required quality threshold at lower latency and cost.

Common Questions

What is Claude Sonnet 5.5?

Claude Sonnet 5.5 is Anthropic’s current fast frontier model for coding, analysis, professional knowledge work, and tool-based agents. Anthropic describes it as the best combination of speed and intelligence.

How much does Claude Sonnet 5.5 cost?

Claude Sonnet 5.5 costs $2 per 1M input tokens and $10 per 1M output tokens. Five-minute cache writes cost $2.50/MTok, one-hour cache writes cost $4/MTok, and cache reads cost $0.20/MTok. Batch input and output receive a 50% discount.

What is the Claude Sonnet 5.5 context window?

Claude Sonnet 5.5 has a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.

Can Claude Sonnet 5.5 output more than 128K tokens?

Anthropic documents a beta maximum of up to 300,000 output tokens through the Message Batches API. Standard requests use the 128K output limit.

What is Claude Sonnet 5.5’s knowledge cutoff?

Anthropic lists June 2026 as both the reliable knowledge cutoff and training-data cutoff.

Does Claude Sonnet 5.5 support images?

Yes. Claude Sonnet 5.5 accepts text and image input and produces text output.

Does Claude Sonnet 5.5 use adaptive thinking?

Yes. Adaptive thinking is enabled by default and the Claude API defaults to high effort. Anthropic recommends lower effort for latency-sensitive workloads and often medium as a starting point for well-specified agentic coding or multistep tool tasks.

What is between_tools in Claude Sonnet 5.5?

between_tools is the model’s lowest thinking mode. It removes up-front thinking before the first response while still allowing short thinking-based progress updates between tool calls. It is supported at low, medium, and high effort.

Can thinking be completely disabled on Claude Sonnet 5.5?

The older thinking disabled setting is not supported and returns an error. Use between_tools when you want to avoid up-front thinking, or adaptive thinking for the normal reasoning path.

Can Claude Sonnet 5.5 force a specific tool call?

No. Anthropic states that tool_choice any and named-tool forcing return a 400 error. Use automatic tool selection and strict tool schemas or structured outputs where appropriate.

What is the minimum cacheable prompt for Claude Sonnet 5.5?

Anthropic documents a 512-token minimum cacheable prompt for Claude Sonnet 5.5, down from 1,024 tokens on Claude Sonnet 5.

How is Claude Sonnet 5.5 different from Claude Opus 5.5?

Both models provide 1M context and 128K standard output, but Sonnet 5.5 is classified as fast and costs $2/$10 per million input/output tokens, while Opus 5.5 has moderate comparative latency and costs $4/$20. Opus is the stronger tier for workloads where Sonnet does not meet the required quality threshold.

How is Claude Sonnet 5.5 different from Claude Sonnet 5?

Sonnet 5.5 keeps the same $2/$10 pricing and tokenizer but changes thinking, tool forcing, preserved thinking blocks, computer use, advisor compatibility, progress streaming, cache minimums, and advanced agent-state controls.

When was Claude Sonnet 5.5 released?

Anthropic lists September 28, 2026 as the release date. Retirement is not scheduled sooner than September 28, 2027.

Should I choose Claude Sonnet 5.5 or Claude Opus 5.5?

Start with your evaluation target. Sonnet 5.5 is the stronger candidate when latency and cost matter and it still meets the required quality threshold. Escalate to Opus 5.5 when difficult coding, reasoning, or long-horizon tasks fail too often on Sonnet.

Model information

Last updated

Specifications, pricing, lifecycle, thinking behavior, tool-use changes, prompt caching, compaction, refusal handling, and migration details on this page are based on Anthropic’s official Claude Sonnet 5.5 overview, What’s New guide, and migration documentation.

Claude Sonnet 5.5 — Pricing, 1M Context, Agentic Coding & Tools | EidoStack