Anthropic model
Claude Sonnet 5.5
Anthropic’s fast frontier model for production coding, tool-heavy agents, and professional knowledge work, combining a 1M-token context window with lower pricing and latency than the Opus tier.
- Context window
- 1M
- tokens
- Max output
- 128K
- tokens
- Input
- $2.00
- per 1M tokens
- Cache read
- $0.20
- per 1M tokens
- Output
- $10.00
- per 1M tokens
01 / Overview
What Claude Sonnet 5.5 Is
Claude Sonnet 5.5 is Anthropic’s current fast frontier model, designed to combine strong intelligence with lower latency and lower token cost than the Opus and Fable tiers.
A Frontier Model Designed to Be Used More Often
Anthropic describes Sonnet 5.5 as its best combination of speed and intelligence.
That positioning matters in production. The highest-capability model is not automatically the best default for every request. A SaaS product may make thousands or millions of model calls, and a modest reduction in per-request cost or latency can compound across the entire workload.
Sonnet 5.5 keeps the same large-scale technical envelope as the current Opus tier: a 1,000,000-token context window, up to 128,000 standard output tokens, text and image input, adaptive thinking, prompt caching, and advanced tool workflows.
The difference is economics and latency. Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, while Anthropic classifies its comparative latency as fast.
Anthropic released Claude Sonnet 5.5 on September 28, 2026 and lists it as Active (latest), with retirement not scheduled sooner than September 28, 2027.
- Model ID:
claude-sonnet-5-5. - 1M-token context window.
- 128K standard maximum output.
- Text and image input with text output.
- Adaptive thinking enabled by default.
- Fast comparative latency.
- $2 input / $10 output per million tokens.
- Provider
- Anthropic
- Family
- Claude Sonnet 5.5
- Model ID
- claude-sonnet-5-5
- Knowledge cutoff
- Jun 2026
- Input modalities
- Text, Image
- Output modality
- Text
- Comparative latency
- Fast
02 / Production fit
When Sonnet 5.5 Makes Sense as the Production Default
Sonnet 5.5 is most compelling when an application needs frontier-level behavior frequently enough that Opus pricing or latency would materially affect unit economics.
Optimize for the Cheapest Model That Reliably Passes the Task
A useful production routing strategy does not ask which model is theoretically strongest. It asks which model clears the required quality threshold for a specific workload.
Sonnet 5.5 can be a strong default for coding assistance, document analysis, customer-facing professional workflows, research synthesis, structured tool use, and multi-step agents where both quality and responsiveness matter.
More difficult tasks can then escalate to Opus or Fable only when evaluation shows that Sonnet fails too often.
This routing pattern can lower total inference cost because the expensive tier is reserved for the minority of requests that actually need it.
The correct decision should come from workload-specific evals. Measure accepted-result rate, retries, human intervention, latency, output length, and end-to-end task cost rather than comparing benchmark scores in isolation.
- 01
Default path
Route routine but meaningful coding, research, analysis, and tool workflows to Sonnet when it consistently meets the acceptance threshold.
- 02
Escalation path
Move unusually difficult or long-horizon failures to Opus or Fable rather than paying premium-model prices on every request.
- 03
Latency-sensitive path
Use lower effort where fast response time matters and evaluation shows that additional reasoning does not improve the outcome.
- 04
Measure the whole task
Compare cost per accepted result, including retries, tool turns, cache activity, output length, and fallbacks.
03 / Agentic coding
Agentic Coding Without Opus-Level Token Pricing
Sonnet 5.5 is particularly relevant for coding agents that need strong repository reasoning and repeated tool use but must remain economical enough for frequent production execution.
Evaluate Completion Efficiency, Not Only Code Quality
An agentic coding workload includes far more than generating a function.
The model may inspect repository structure, read files, search symbols, plan a change, call tools, edit several files, run tests, investigate failures, and repeat until the implementation is complete.
For these tasks, the useful comparison is trajectory efficiency:
- How often does the agent finish successfully?
- How many tool calls does it need?
- How much context does it accumulate?
- How often does it repeat work?
- How many tokens are consumed before the tests pass?
- When does a human need to intervene?
Anthropic’s guidance for Sonnet 5.5 recommends starting well-specified agentic coding and multistep tool workloads around medium effort and moving toward high for harder or longer tasks.
That is economically important. A lower-cost model running a shorter, cleaner trajectory can outperform a stronger model on total cost even when the stronger model has a higher one-shot success rate.
- 01
Feature implementation
Navigate a repository, modify multiple files, call tools, run tests, and keep iterating until the requested feature is complete.
- 02
Debugging
Trace failures through code, logs, configuration, and tool output while keeping the original problem in working context.
- 03
Code review
Inspect implementation details and return actionable findings without making every review request an Opus-tier call.
- 04
High-volume engineering automation
Run repeated coding and maintenance jobs where per-task latency and inference cost compound quickly.
04 / Pricing
Claude Sonnet 5.5 Pricing
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, with the same base pricing as Claude Sonnet 5.
The Economics Are Designed for Frequent Production Use
Prompt caching can reduce repeated-context cost further.
Anthropic lists 5-minute cache writes at $2.50 per million tokens, 1-hour cache writes at $4 per million, and cache reads at $0.20 per million.
The model also supports Batch API processing with a 50% discount on standard input and output. That can be useful for offline evaluations, bulk document processing, code analysis, or other workloads that do not need interactive latency.
Sonnet 5.5 is half the standard token price of Opus 5.5: $2/$10 versus $4/$20. Fable 5.1 is substantially more expensive at $10/$50.
This is why model routing matters. If Sonnet reaches the same application acceptance threshold as a premium tier on most requests, the cost difference can be significant at scale.
- Input: $2.00 per 1M tokens.
- Output: $10.00 per 1M tokens.
- 5-minute cache write: $2.50 per 1M tokens.
- 1-hour cache write: $4.00 per 1M tokens.
- Cache read: $0.20 per 1M tokens.
- Batch API: 50% discount on input and output.
1M tokens · USD
- Input
- $2.00
- 5m cache write
- $2.50
- 1h cache write
- $4.00
- Cache read
- $0.20
- Output
- $10.00
Example: 20K input + 4K output
- Input cost
- $0.0400
- Output cost
- $0.0400
- Estimated uncached total
- $0.0800
05 / Context
1M Context and 128K Standard Output
Claude Sonnet 5.5 provides a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.
Large Enough for Repository, Research, and Agent State
A million-token context can hold a large repository slice, long documents, screenshots, conversation history, tool results, system instructions, retrieved evidence, and intermediate agent state.
This makes Sonnet 5.5 suitable for workloads where the application wants a large working set without immediately escalating to a more expensive model tier.
Anthropic also documents up to 300,000 output tokens through the Message Batches API in beta. That higher limit is specific to batch processing; standard requests retain the 128K output ceiling.
Prompt caching becomes especially relevant with large contexts. Sonnet 5.5 lowers the minimum cacheable prompt to 512 tokens, compared with 1,024 tokens on Sonnet 5.
A larger context window still does not remove the need for retrieval and state management. Sending irrelevant history consumes money and can reduce signal quality even when it technically fits.
Context window
1,000,000
Standard max output
128,000
Message Batches can expose up to 300K output tokens in beta. Prompt caching supports a 512-token minimum cacheable prompt.
06 / Thinking
Adaptive Thinking With a `between_tools` Low-Compute Path
Claude Sonnet 5.5 uses adaptive thinking by default, but unlike the always-on Opus 5.5 contract, it also provides between_tools as a lower-compute mode that removes up-front thinking while preserving short reasoning updates during tool workflows.
Thinking Can Be Tuned for Latency-Sensitive Work
On the Claude API, the default effort level is high.
For well-specified agentic coding and multistep tool use, Anthropic recommends starting around medium and moving to high for more difficult or longer tasks. For chat and latency-sensitive workloads, medium or low can be a better starting point.
Sonnet 5.5 also introduces a specific alternative to disabled thinking: thinking: {"type": "between_tools"}.
With between_tools, the model does not perform up-front thinking before the first response, but can still generate progress-oriented thinking blocks between tool calls. Anthropic describes it as the lowest thinking setting on the model.
between_tools works at low, medium, and high effort. At xhigh or max, the request must use adaptive thinking instead.
The older thinking: disabled mode returns a 400 error, as does manual budget_tokens thinking.
Latency-sensitive interaction
Start with low or medium effort, or use between_tools when the workflow uses tools but does not need up-front reasoning.
Hard agentic task
Use adaptive thinking at high or above when deeper reasoning materially improves completion quality.
07 / Tools & state
Dynamic Tools, Compaction, and Long-Running State
Sonnet 5.5 adds several controls that are particularly useful for persistent agents: mid-conversation system messages, mid-conversation tool changes, on-demand compaction, and inline tool definitions.
Modify the Agent Without Rebuilding the Conversation
Anthropic supports mid-conversation system messages on Sonnet 5.5. This lets an application update instructions during a session without editing earlier turns.
Mid-conversation tool changes are available in beta. An application can change the available tool set as a task progresses instead of exposing every tool from the beginning.
Anthropic also documents compact on demand in beta. The API can return a signed compaction block that summarizes the prior conversation so the application can replace older raw turns while preserving the working state needed to continue.
Another beta feature allows tool definitions to be introduced inside a mid-conversation system message. This can add a tool, modify its schema, or move to a newer server-tool version without rewriting the top-level tool list and losing earlier prompt-cache continuity.
These controls make Sonnet 5.5 especially relevant for agents whose state, permissions, and available actions evolve over time.
- Supported
Mid-conversation system messages
Update instructions during a session without editing the earlier conversation history.
- Supported
Mid-conversation tool changes
Beta support lets applications add or remove tools as the agent progresses.
- Supported
Compaction on demand
Beta compaction can replace older conversation history with a signed summarized state.
- Supported
Inline tool definitions
Beta system-message blocks can add or change full tool definitions without rebuilding the top-level tool list.
- Supported
Prompt caching
A 512-token minimum makes shorter reusable prefixes eligible for caching.
- Supported
Computer use
On the Claude API and Google Cloud, Sonnet 5.5 uses the computer_toolset_20260801 interface.
08 / Refusals
Refusals Are a Structured Application Outcome
Claude Sonnet 5.5 can decline requests with stop_reason: "refusal" and structured stop_details, which lets applications handle different safety outcomes explicitly instead of treating them as generic failures.
Production Systems Should Route Refusals Deliberately
Anthropic documents five safeguard categories for Sonnet 5.5:
cyberbiofrontier_llmreasoning_extractiongeneral_harms
A refusal still returns HTTP 200. The application should inspect the stop reason and metadata rather than assume that a successful HTTP response always contains an ordinary completion.
Anthropic also provides beta server-side fallback behavior for some refusal classes. For example, configured fallback can retry certain cyber or frontier_llm declines on Claude Sonnet 5.
This means refusal handling belongs in production orchestration. Track refusal rate by workload, distinguish policy outcomes from ordinary model failures, and decide whether a safe fallback, user-facing explanation, or no retry is appropriate.
- 01
Inspect stop_reason
Treat refusal as a distinct successful API response shape rather than a transport failure.
- 02
Read stop_details
Use the documented safeguard category to understand what class of request was declined.
- 03
Configure safe fallback
Use provider-supported or application-level fallbacks only for categories where retry behavior is appropriate.
- 04
Measure refusal rate
Include refusals in production evaluation because they affect task completion, routing economics, and user experience.
09 / Sonnet 5 → 5.5
Migrating From Claude Sonnet 5 to Sonnet 5.5
Sonnet 5.5 keeps Sonnet 5 pricing and tokenization, but several thinking, tool, computer-use, and conversation-state behaviors change enough to require integration testing.
Same Economics Does Not Mean Drop-In Compatibility
The tokenizer is the same as Sonnet 5, so identical text should produce the same token counts. Standard prices also remain $2 input and $10 output per million tokens.
The breaking changes are mostly behavioral and orchestration-related.
First, thinking: disabled is replaced by between_tools when an application wants to remove up-front thinking.
Second, forced tool use is no longer supported. tool_choice: any and named-tool forcing return errors.
Third, thinking blocks are bound to the model and conversation. Editing the system prompt, tools, or earlier messages can invalidate preserved Sonnet 5.5 thinking blocks, so append-only conversation design is safer.
Fourth, computer use on the Claude API and Google Cloud moves from the older computer_20251124 tool to computer_toolset_20260801.
Fifth, some advisor-model pairings that worked with Sonnet 5 are not accepted by Sonnet 5.5.
Finally, progress text between tool calls may now arrive in thinking blocks. Applications that stream agent status to users should test thinking.display or between_tools behavior explicitly.
- 01
Replace disabled thinking
Use between_tools at supported effort levels instead of thinking disabled.
- 02
Remove forced tool selection
Move any-tool or named-tool forcing to automatic selection plus strict schemas or structured output patterns.
- 03
Keep histories append-only
Avoid editing earlier prompts, tools, or messages when preserved thinking blocks are replayed.
- 04
Update agent integrations
Retest computer use, advisor pairings, progress streaming, dynamic tools, and conversation-state handling.
10 / Evaluation
Claude Sonnet 5.5 Strengths and Limitations
Sonnet 5.5 is a strong production-oriented frontier model because it combines 1M context and advanced agent controls with fast comparative latency and half the base token price of Opus 5.5.
Strengths
Strong price-performance
$2/$10 standard pricing makes frontier-level coding, analysis, and tool workflows substantially cheaper than current Opus or Fable tiers.
Fast frontier latency
Anthropic classifies Sonnet 5.5 as fast, making it suitable for interactive and high-volume production paths.
1M context and 128K output
Large working-state capacity supports repositories, long documents, images, agent history, and substantial generated artifacts.
Advanced agent-state controls
Mid-conversation instructions, dynamic tools, compaction, prompt caching, and between_tools reasoning support sophisticated long-running workflows.
What to consider
Not Anthropic’s maximum capability tier
Difficult long-horizon tasks may still justify escalation to Opus 5.5 or Fable 5.1 when Sonnet fails the required evaluation threshold.
Forced tool use is unavailable
Applications that depend on guaranteed named-tool invocation need a different orchestration pattern.
Thinking state needs careful handling
Preserved thinking blocks are tied to model and conversation state, so editing historical context can create compatibility errors.
Several advanced controls are beta
Per-message effort, dynamic tool changes, compaction on demand, and inline tool definitions may evolve and should be isolated behind stable application abstractions.
Optimize the production default
Evaluate Claude Sonnet 5.5 on your real workload
Run the same coding, tool-use, research, document, and long-context tasks across Sonnet and Opus, then compare accepted-result quality, latency, token usage, cache efficiency, and total cost per completed task.
Start FreeFor many production systems, the useful question is not whether Sonnet is the strongest model available, but whether it reaches the required quality threshold at lower latency and cost.
Common Questions
What is Claude Sonnet 5.5?
Claude Sonnet 5.5 is Anthropic’s current fast frontier model for coding, analysis, professional knowledge work, and tool-based agents. Anthropic describes it as the best combination of speed and intelligence.
How much does Claude Sonnet 5.5 cost?
Claude Sonnet 5.5 costs $2 per 1M input tokens and $10 per 1M output tokens. Five-minute cache writes cost $2.50/MTok, one-hour cache writes cost $4/MTok, and cache reads cost $0.20/MTok. Batch input and output receive a 50% discount.
What is the Claude Sonnet 5.5 context window?
Claude Sonnet 5.5 has a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.
Can Claude Sonnet 5.5 output more than 128K tokens?
Anthropic documents a beta maximum of up to 300,000 output tokens through the Message Batches API. Standard requests use the 128K output limit.
What is Claude Sonnet 5.5’s knowledge cutoff?
Anthropic lists June 2026 as both the reliable knowledge cutoff and training-data cutoff.
Does Claude Sonnet 5.5 support images?
Yes. Claude Sonnet 5.5 accepts text and image input and produces text output.
Does Claude Sonnet 5.5 use adaptive thinking?
Yes. Adaptive thinking is enabled by default and the Claude API defaults to high effort. Anthropic recommends lower effort for latency-sensitive workloads and often medium as a starting point for well-specified agentic coding or multistep tool tasks.
What is between_tools in Claude Sonnet 5.5?
between_tools is the model’s lowest thinking mode. It removes up-front thinking before the first response while still allowing short thinking-based progress updates between tool calls. It is supported at low, medium, and high effort.
Can thinking be completely disabled on Claude Sonnet 5.5?
The older thinking disabled setting is not supported and returns an error. Use between_tools when you want to avoid up-front thinking, or adaptive thinking for the normal reasoning path.
Can Claude Sonnet 5.5 force a specific tool call?
No. Anthropic states that tool_choice any and named-tool forcing return a 400 error. Use automatic tool selection and strict tool schemas or structured outputs where appropriate.
What is the minimum cacheable prompt for Claude Sonnet 5.5?
Anthropic documents a 512-token minimum cacheable prompt for Claude Sonnet 5.5, down from 1,024 tokens on Claude Sonnet 5.
How is Claude Sonnet 5.5 different from Claude Opus 5.5?
Both models provide 1M context and 128K standard output, but Sonnet 5.5 is classified as fast and costs $2/$10 per million input/output tokens, while Opus 5.5 has moderate comparative latency and costs $4/$20. Opus is the stronger tier for workloads where Sonnet does not meet the required quality threshold.
How is Claude Sonnet 5.5 different from Claude Sonnet 5?
Sonnet 5.5 keeps the same $2/$10 pricing and tokenizer but changes thinking, tool forcing, preserved thinking blocks, computer use, advisor compatibility, progress streaming, cache minimums, and advanced agent-state controls.
When was Claude Sonnet 5.5 released?
Anthropic lists September 28, 2026 as the release date. Retirement is not scheduled sooner than September 28, 2027.
Should I choose Claude Sonnet 5.5 or Claude Opus 5.5?
Start with your evaluation target. Sonnet 5.5 is the stronger candidate when latency and cost matter and it still meets the required quality threshold. Escalate to Opus 5.5 when difficult coding, reasoning, or long-horizon tasks fail too often on Sonnet.
Model information
Last updated
Specifications, pricing, lifecycle, thinking behavior, tool-use changes, prompt caching, compaction, refusal handling, and migration details on this page are based on Anthropic’s official Claude Sonnet 5.5 overview, What’s New guide, and migration documentation.