Anthropic legacy frontier model
Claude Sonnet 5
The first Claude 5 Sonnet model: a fast frontier baseline for coding, analysis, and agentic workflows with a 1M-token context window, 128K output, adaptive thinking on by default, and lower pricing than Sonnet 4.6.
- Context window
- 1M
- tokens
- Max output
- 128K
- tokens
- Input
- $2.00
- per 1M tokens
- Cache read
- $0.20
- per 1M tokens
- Output
- $10.00
- per 1M tokens
01 / Overview
What Claude Sonnet 5 Is
Claude Sonnet 5 is Anthropic’s first Claude 5 Sonnet model, launched as the next generation of the Sonnet family for coding, analysis, agentic work, and other production workloads that need a strong balance of intelligence, speed, and cost.
The First Sonnet of the Claude 5 Generation
Anthropic released Claude Sonnet 5 on June 30, 2026.
At launch, it represented a capability upgrade over Sonnet 4.6 while lowering the standard per-token price from $3/$15 to $2/$10 per million input/output tokens. Anthropic later made the $2/$10 rate the standard price.
The model provides a 1,000,000-token context window and up to 128,000 output tokens in standard requests. It accepts text and images and produces text.
Sonnet 5 also marked an important API transition. Adaptive thinking became active by default, manual budget_tokens thinking was removed, non-default sampling parameters became invalid, and a new tokenizer changed token counts for the same text.
The model is now listed as Active (legacy). Anthropic recommends migrating to Claude Sonnet 5.5 for improved performance.
- Model ID:
claude-sonnet-5. - Released June 30, 2026.
- Active legacy model.
- 1M context and 128K standard output.
- Text and image input with text output.
- Adaptive thinking on by default.
- Fast Sonnet-tier production positioning.
- Provider
- Anthropic
- Family
- Claude Sonnet 5
- Model ID
- claude-sonnet-5
- Knowledge cutoff
- Jan 2026
- Input modalities
- Text, Image
- Output modality
- Text
- Lifecycle
- Active · legacy
02 / Coding & agents
A Major Sonnet Upgrade for Coding and Agentic Work
Anthropic identifies coding and agentic tasks as the largest capability gains of Sonnet 5 over Sonnet 4.6.
A Lower-Cost Alternative to Moving Every Hard Task to Opus
Sonnet 5 created a useful middle path for production systems.
A workload that exceeded Sonnet 4.6 capability no longer had to jump immediately to an Opus-tier model. Sonnet 5 increased capability while simultaneously lowering the standard token rate.
That is especially relevant for coding agents because the cost of one task is not just one completion. An agent may inspect repository files, search symbols, call tools, edit code, run tests, investigate failures, and repeat several times before the task is accepted.
For that reason, evaluate the complete trajectory:
- task completion rate;
- number of tool calls;
- retries and repeated work;
- test or validation success;
- human interventions;
- thinking and output consumption;
- latency to accepted result;
- total task cost.
Anthropic’s prompting guidance recommends xhigh effort for the hardest coding and agentic workloads, while high is the default general setting.
- 01
Repository-scale implementation
Navigate source files, dependencies, tests, and tooling across a multi-step implementation rather than generating an isolated snippet.
- 02
Debugging loops
Inspect symptoms, call tools, revise hypotheses, modify code, and repeat validation until the root cause is resolved.
- 03
Code review
Analyze implementation details and surface actionable defects without requiring an Opus-tier call for every review.
- 04
High-volume engineering automation
Run repeated maintenance, migration, refactoring, and analysis tasks where per-task economics matter at scale.
03 / Pricing
Claude Sonnet 5 Pricing
Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens, with prompt caching and discounted Batch API processing.
Lower Per-Token Pricing Than Sonnet 4.6
Sonnet 4.6 costs $3 input and $15 output per million tokens. Sonnet 5 lowered those rates to $2 and $10 while increasing capability.
However, the economic improvement is not a simple one-third reduction for identical text because Sonnet 5 introduced a new tokenizer that produces roughly 30% more tokens than Sonnet 4.6 for the same content.
Prompt caching helps reduce repeated-input cost. Anthropic lists 5-minute cache writes at $2.50 per million tokens, 1-hour cache writes at $4 per million, and cache reads at $0.20 per million.
Batch API input and output receive a 50% discount.
For production evaluation, measure actual usage rather than multiplying an old Sonnet 4.6 token count by the new price.
- Input: $2.00 per 1M tokens.
- Output: $10.00 per 1M tokens.
- 5-minute cache write: $2.50 per 1M tokens.
- 1-hour cache write: $4.00 per 1M tokens.
- Cache read: $0.20 per 1M tokens.
- Batch API: 50% discount on input and output.
1M tokens · USD
- Input
- $2.00
- 5m cache write
- $2.50
- 1h cache write
- $4.00
- Cache read
- $0.20
- Output
- $10.00
Example: 20K input + 4K output
- Input cost
- $0.0400
- Output cost
- $0.0400
- Estimated uncached total
- $0.0800
04 / Context
A 1M Context Window With 128K Standard Output
Claude Sonnet 5 provides a 1,000,000-token context window by default and supports up to 128,000 output tokens in standard requests.
Large Working State Without Moving to Opus
The million-token window gives a Sonnet-tier model room for repositories, long reports, screenshots, conversation history, tool results, system instructions, retrieved evidence, and intermediate agent state.
That can be useful for products that need large working context but cannot justify Opus-tier economics for every request.
Anthropic also documents a beta path for up to 300,000 output tokens through the Message Batches API.
Prompt caching supports repeated large prefixes such as system prompts, repositories, policies, or long reference documents. On Sonnet 5, the minimum cacheable prompt is 1,024 tokens.
A large context ceiling still does not mean every request should fill the window. Relevance filtering, caching, retrieval, and explicit context management can reduce both cost and distraction.
Context window
1,000,000
Standard max output
128,000
Message Batches can expose up to 300K output tokens in beta. Sonnet 5 uses a 1,024-token minimum cacheable prompt.
05 / Thinking & effort
Adaptive Thinking Is On by Default—but Can Be Disabled
Claude Sonnet 5 made adaptive thinking the default Sonnet behavior while still allowing applications to turn thinking off completely when latency or workload simplicity makes reasoning unnecessary.
A Different Contract From Both Sonnet 4.6 and Sonnet 5.5
On Sonnet 4.6, a request that omitted thinking ran without thinking.
On Sonnet 5, the same request runs with adaptive thinking by default. The model decides when and how much to reason, while the effort parameter controls how much computational work it should invest.
Manual extended thinking is no longer accepted. Requests using thinking: {"type": "enabled", "budget_tokens": N} return a 400 error.
Unlike Sonnet 5.5, however, Sonnet 5 still accepts thinking: {"type": "disabled"}. That makes Sonnet 5 a distinctive compatibility point for applications that want Claude 5 capability but still need a fully non-thinking path for certain requests.
Anthropic supports five effort levels: low, medium, high, xhigh, and max. high is the default.
Latency-sensitive request
Use low effort or disable thinking when evaluations show that reasoning adds latency without improving the accepted result.
Hard coding or agentic task
Anthropic recommends xhigh for the most difficult coding and agentic workloads, with max reserved for cases where evals justify unrestricted effort.
06 / Tokenizer
Sonnet 5 Introduced a New Tokenizer
Claude Sonnet 5 changed tokenization enough that the same text produces approximately 30% more tokens than on Sonnet 4.6, depending on content.
Lower Token Prices Do Not Translate Directly Into the Same Percentage of Savings
The API request and response shapes did not change because of the tokenizer, but token-based assumptions did.
Anything that depends on token counts needs to be recalibrated:
- prompt-size forecasts;
- monthly cost estimates;
- context-window thresholds;
- compaction or trimming rules;
max_tokensvalues;- rate-limit planning;
- benchmark comparisons.
The nominal context window remains 1M tokens, but if each token covers less text on average, the same 1M-token window can hold less equivalent source text than Sonnet 4.6.
Similarly, an output limit tuned closely around Sonnet 4.6 responses may truncate equivalent Sonnet 5 output.
Anthropic recommends recounting prompts with the token-counting API instead of carrying forward old estimates.
- 01
Recount prompts
Measure real production payloads under Sonnet 5 instead of reusing Sonnet 4.6 token counts.
- 02
Revisit max_tokens
Equivalent text can require more tokens, so output ceilings tuned for earlier Sonnet models may truncate responses.
- 03
Recheck context thresholds
Retrieval, trimming, and context-management logic based on old token ratios can trigger at the wrong point.
- 04
Recalculate effective savings
The $2/$10 rate is lower than Sonnet 4.6, but the tokenizer means equivalent-request savings are workload-dependent.
07 / Tools
A More Permissive Tool Contract Than Sonnet 5.5
Claude Sonnet 5 supports the same general tool and platform feature set as Sonnet 4.6, except Priority Tier, and it retains forced tool-choice patterns that Sonnet 5.5 later removes.
Useful for Existing Agents That Depend on Explicit Tool Forcing
Sonnet 5 accepts tool_choice modes that force any tool or a specific named tool.
That can matter in deterministic application flows where the product expects the model to call a tool rather than return free-form text.
Sonnet 5.5 changes this contract: forced any or named-tool choice returns a 400 error and applications must move toward automatic tool selection with strict schemas and stronger prompt instructions.
Sonnet 5 also predates several Sonnet 5.5-specific agent controls such as per-message effort, mid-conversation system messages, mid-conversation tool changes, and the lower 512-token prompt-cache minimum.
Assistant-message prefilling is not supported on Sonnet 5, as was already true for Sonnet 4.6.
Priority Tier is also unavailable on Sonnet 5.
- Supported
Tool use
Supports server-side and client-defined tool workflows for agentic applications.
- Supported
Forced tool choice
Can force any tool or a named tool, unlike Sonnet 5.5.
- Supported
Prompt caching
Cache stable prompt prefixes with a 1,024-token minimum cacheable prompt.
- Supported
Image input
Combine supported images with text and long-context workflows.
- Not listed
Assistant prefill
A prefilled final assistant turn returns an error.
- Not listed
Priority Tier
Anthropic explicitly excludes Priority Tier from the Sonnet 5 platform feature set.
08 / Safeguards
The First Sonnet Model With Real-Time Cybersecurity Safeguards
Claude Sonnet 5 introduced real-time cybersecurity safeguards to the Sonnet tier, making refusal handling a production integration concern for security-sensitive workloads.
Refusal Is a Model Outcome, Not an HTTP Error
Anthropic documents that prohibited or high-risk cybersecurity requests can be declined by the model.
A refusal is returned as a successful HTTP 200 response with stop_reason: "refusal" rather than as a transport error.
That means production applications should inspect the response stop reason explicitly. A generic handler that treats every HTTP 200 as an ordinary completion may misclassify a safety outcome as usable model output.
This is particularly relevant for legitimate security products because benign defensive work can overlap with risky technical domains. Anthropic provides a Cyber Verification Program for eligible legitimate security work that needs reduced restrictions.
Sonnet 5.5 later expands safeguard categories and fallback behavior, but Sonnet 5 remains the first Sonnet baseline where real-time cyber refusal handling becomes part of the API integration story.
- 01
Inspect stop_reason
Treat refusal as a distinct model outcome even though the API request itself succeeds with HTTP 200.
- 02
Track refusal rates
Include policy refusals in production completion metrics because they affect accepted-result rates and fallback behavior.
- 03
Separate policy from model failure
Do not classify a safeguard refusal as a parsing, transport, or ordinary quality failure.
- 04
Validate security workflows
Legitimate cybersecurity products should evaluate safeguard behavior and provider verification options before production rollout.
09 / Sonnet 5 → 5.5
Migrating From Claude Sonnet 5 to Sonnet 5.5
Sonnet 5.5 keeps the same $2/$10 pricing and the same tokenizer as Sonnet 5, but changes thinking, forced tool use, conversation-state behavior, cache minimums, progress streaming, and several agent integrations.
The Economics Stay Stable While the Orchestration Contract Changes
The easiest part of the migration is token accounting: Sonnet 5.5 uses the same tokenizer and standard prices as Sonnet 5.
The harder part is API behavior.
Sonnet 5 can turn thinking fully off with thinking: disabled. Sonnet 5.5 rejects that setting and introduces between_tools as the lowest-thinking alternative.
Sonnet 5 accepts forced tool choice. Sonnet 5.5 rejects tool_choice: any and named-tool forcing.
The prompt-cache minimum also falls from 1,024 tokens on Sonnet 5 to 512 tokens on Sonnet 5.5.
Sonnet 5.5 adds per-message effort, mid-conversation system messages, and mid-conversation tool changes. It also binds preserved thinking blocks more tightly to model and conversation state, so append-only histories become safer.
Progress text changes shape too. On Sonnet 5, text written between tool calls is returned as ordinary text blocks. On Sonnet 5.5, longer progress notes can arrive as thinking blocks.
Computer-use and advisor-tool compatibility also change.
- 01
Replace disabled thinking
If the application disables thinking on Sonnet 5, migrate that path to between_tools or adaptive thinking on Sonnet 5.5.
- 02
Remove forced tool selection
Replace any-tool or named-tool forcing with automatic selection, strict tool schemas, and explicit prompting.
- 03
Review conversation state
Keep preserved thinking histories append-only and test progress streaming, computer use, and advisor integrations.
- 04
Take advantage of newer controls
Re-evaluate caching, per-message effort, dynamic tools, and mid-conversation instructions after the migration.
10 / Evaluation
Claude Sonnet 5 Strengths and Limitations
Claude Sonnet 5 remains a useful production and migration baseline because it combines current-scale context and low Sonnet pricing with a more permissive thinking and tool contract than Sonnet 5.5.
Strengths
Strong coding and agentic baseline
Anthropic identifies coding and agentic tasks as the largest capability gains over Sonnet 4.6.
Low Sonnet-tier pricing
$2/$10 standard pricing supports frequent production use without Opus-tier token costs.
Flexible thinking behavior
Adaptive thinking runs by default but can still be completely disabled for latency-sensitive or simple workloads.
Permissive forced-tool support
Existing agents can force any or named tools, a behavior removed in Sonnet 5.5.
What to consider
Legacy model
Anthropic recommends migrating to Claude Sonnet 5.5 for improved performance.
New tokenizer complicates Sonnet 4.6 comparisons
The same text uses roughly 30% more tokens than Sonnet 4.6, so historical token and cost assumptions need recalibration.
Fewer dynamic agent controls than Sonnet 5.5
Per-message effort, mid-conversation system messages, dynamic tool changes, and a lower cache threshold arrive in the successor.
No Priority Tier
Applications that depend on Priority Tier need a different model or deployment strategy.
Preserve the first Claude 5 Sonnet baseline
Compare Claude Sonnet 5 with Sonnet 5.5 on your own workload
Replay coding, analysis, tool-use, long-context, and latency-sensitive requests to measure how the newer Sonnet changes quality, thinking behavior, tool orchestration, caching, and total cost.
Start FreeClaude Sonnet 5 remains active as a legacy model, while Anthropic recommends migrating to Claude Sonnet 5.5 for improved performance.
Common Questions
What is Claude Sonnet 5?
Claude Sonnet 5 is Anthropic’s first Claude 5 Sonnet model, released on June 30, 2026 as a capability upgrade over Sonnet 4.6 with lower standard pricing and stronger coding and agentic performance.
How much does Claude Sonnet 5 cost?
Claude Sonnet 5 costs $2 per 1M input tokens and $10 per 1M output tokens. Five-minute cache writes cost $2.50/MTok, one-hour cache writes cost $4/MTok, and cache reads cost $0.20/MTok. Batch input and output receive a 50% discount.
What is the Claude Sonnet 5 context window?
Claude Sonnet 5 has a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.
Can Claude Sonnet 5 output more than 128K tokens?
Anthropic documents a beta maximum of up to 300,000 output tokens through the Message Batches API. Standard requests use the 128K output limit.
What is Claude Sonnet 5’s knowledge cutoff?
Anthropic lists January 2026 as both the reliable knowledge cutoff and the training-data cutoff.
Does Claude Sonnet 5 support images?
Yes. Claude Sonnet 5 accepts text and image input and produces text output.
Does Claude Sonnet 5 use adaptive thinking?
Yes. Adaptive thinking is on by default, with high as the default effort level. The model supports low, medium, high, xhigh, and max effort.
Can thinking be disabled on Claude Sonnet 5?
Yes. Claude Sonnet 5 accepts thinking type disabled. This differs from Claude Sonnet 5.5, where the disabled setting is rejected and between_tools is used to remove up-front thinking.
Does Claude Sonnet 5 support budget_tokens?
No. Manual extended thinking with thinking type enabled and budget_tokens was removed in Sonnet 5 and returns a 400 error. Use adaptive thinking with effort instead.
Why does Claude Sonnet 5 use more tokens than Sonnet 4.6?
Sonnet 5 introduced a new tokenizer. Anthropic states that the same text produces approximately 30% more tokens than on Sonnet 4.6, depending on content.
Can Claude Sonnet 5 force a tool call?
Yes. Sonnet 5 accepts forced any-tool and named-tool selection. Sonnet 5.5 removes that behavior, so forced-tool integrations require migration changes.
Does Claude Sonnet 5 support Priority Tier?
No. Anthropic states that Sonnet 5 supports the same tools and platform features as Sonnet 4.6 except Priority Tier.
What is the minimum cacheable prompt for Claude Sonnet 5?
Anthropic documents a 1,024-token minimum cacheable prompt for Sonnet 5. Sonnet 5.5 lowers that minimum to 512 tokens.
How is Claude Sonnet 5 different from Claude Sonnet 5.5?
Both use the same tokenizer, $2/$10 pricing, 1M context, and 128K standard output. Sonnet 5 allows thinking disabled and forced tool choice, while Sonnet 5.5 replaces disabled thinking with between_tools, removes forced tool selection, lowers the cache minimum, and adds newer dynamic agent controls.
Is Claude Sonnet 5 still available?
Yes. Anthropic lists Claude Sonnet 5 as Active (legacy). It was released June 30, 2026, with retirement not scheduled sooner than June 30, 2027.
Should I use Claude Sonnet 5 for a new application?
For new development, evaluate Claude Sonnet 5.5 first because Anthropic recommends the newer model. Sonnet 5 remains useful for existing integrations, forced-tool workflows, disabled-thinking compatibility, regression testing, and migration baselines.
Model information
Last updated
Specifications, pricing, lifecycle, thinking behavior, tokenizer changes, tool compatibility, cybersecurity safeguards, and migration details on this page are based on Anthropic’s official Claude Sonnet 5 overview, What’s New guide, prompting guidance, effort documentation, and Sonnet 5.5 migration guide.