Anthropic legacy frontier model

Claude Sonnet 5

The first Claude 5 Sonnet model: a fast frontier baseline for coding, analysis, and agentic workflows with a 1M-token context window, 128K output, adaptive thinking on by default, and lower pricing than Sonnet 4.6.

Context window
1M
tokens
Max output
128K
tokens
Input
$2.00
per 1M tokens
Cache read
$0.20
per 1M tokens
Output
$10.00
per 1M tokens

01 / Overview

What Claude Sonnet 5 Is

Claude Sonnet 5 is Anthropic’s first Claude 5 Sonnet model, launched as the next generation of the Sonnet family for coding, analysis, agentic work, and other production workloads that need a strong balance of intelligence, speed, and cost.

The First Sonnet of the Claude 5 Generation

Anthropic released Claude Sonnet 5 on June 30, 2026.

At launch, it represented a capability upgrade over Sonnet 4.6 while lowering the standard per-token price from $3/$15 to $2/$10 per million input/output tokens. Anthropic later made the $2/$10 rate the standard price.

The model provides a 1,000,000-token context window and up to 128,000 output tokens in standard requests. It accepts text and images and produces text.

Sonnet 5 also marked an important API transition. Adaptive thinking became active by default, manual budget_tokens thinking was removed, non-default sampling parameters became invalid, and a new tokenizer changed token counts for the same text.

The model is now listed as Active (legacy). Anthropic recommends migrating to Claude Sonnet 5.5 for improved performance.

  • Model ID: claude-sonnet-5.
  • Released June 30, 2026.
  • Active legacy model.
  • 1M context and 128K standard output.
  • Text and image input with text output.
  • Adaptive thinking on by default.
  • Fast Sonnet-tier production positioning.
Model profile
Provider
Anthropic
Family
Claude Sonnet 5
Model ID
claude-sonnet-5
Knowledge cutoff
Jan 2026
Input modalities
Text, Image
Output modality
Text
Lifecycle
Active · legacy

02 / Coding & agents

A Major Sonnet Upgrade for Coding and Agentic Work

Anthropic identifies coding and agentic tasks as the largest capability gains of Sonnet 5 over Sonnet 4.6.

A Lower-Cost Alternative to Moving Every Hard Task to Opus

Sonnet 5 created a useful middle path for production systems.

A workload that exceeded Sonnet 4.6 capability no longer had to jump immediately to an Opus-tier model. Sonnet 5 increased capability while simultaneously lowering the standard token rate.

That is especially relevant for coding agents because the cost of one task is not just one completion. An agent may inspect repository files, search symbols, call tools, edit code, run tests, investigate failures, and repeat several times before the task is accepted.

For that reason, evaluate the complete trajectory:

  • task completion rate;
  • number of tool calls;
  • retries and repeated work;
  • test or validation success;
  • human interventions;
  • thinking and output consumption;
  • latency to accepted result;
  • total task cost.

Anthropic’s prompting guidance recommends xhigh effort for the hardest coding and agentic workloads, while high is the default general setting.

Agentic production workloads
  1. 01

    Repository-scale implementation

    Navigate source files, dependencies, tests, and tooling across a multi-step implementation rather than generating an isolated snippet.

  2. 02

    Debugging loops

    Inspect symptoms, call tools, revise hypotheses, modify code, and repeat validation until the root cause is resolved.

  3. 03

    Code review

    Analyze implementation details and surface actionable defects without requiring an Opus-tier call for every review.

  4. 04

    High-volume engineering automation

    Run repeated maintenance, migration, refactoring, and analysis tasks where per-task economics matter at scale.

03 / Pricing

Claude Sonnet 5 Pricing

Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens, with prompt caching and discounted Batch API processing.

Lower Per-Token Pricing Than Sonnet 4.6

Sonnet 4.6 costs $3 input and $15 output per million tokens. Sonnet 5 lowered those rates to $2 and $10 while increasing capability.

However, the economic improvement is not a simple one-third reduction for identical text because Sonnet 5 introduced a new tokenizer that produces roughly 30% more tokens than Sonnet 4.6 for the same content.

Prompt caching helps reduce repeated-input cost. Anthropic lists 5-minute cache writes at $2.50 per million tokens, 1-hour cache writes at $4 per million, and cache reads at $0.20 per million.

Batch API input and output receive a 50% discount.

For production evaluation, measure actual usage rather than multiplying an old Sonnet 4.6 token count by the new price.

  • Input: $2.00 per 1M tokens.
  • Output: $10.00 per 1M tokens.
  • 5-minute cache write: $2.50 per 1M tokens.
  • 1-hour cache write: $4.00 per 1M tokens.
  • Cache read: $0.20 per 1M tokens.
  • Batch API: 50% discount on input and output.
Token and cache pricing

1M tokens · USD

Input
$2.00
5m cache write
$2.50
1h cache write
$4.00
Cache read
$0.20
Output
$10.00

Example: 20K input + 4K output

Input cost
$0.0400
Output cost
$0.0400
Estimated uncached total
$0.0800

04 / Context

A 1M Context Window With 128K Standard Output

Claude Sonnet 5 provides a 1,000,000-token context window by default and supports up to 128,000 output tokens in standard requests.

Large Working State Without Moving to Opus

The million-token window gives a Sonnet-tier model room for repositories, long reports, screenshots, conversation history, tool results, system instructions, retrieved evidence, and intermediate agent state.

That can be useful for products that need large working context but cannot justify Opus-tier economics for every request.

Anthropic also documents a beta path for up to 300,000 output tokens through the Message Batches API.

Prompt caching supports repeated large prefixes such as system prompts, repositories, policies, or long reference documents. On Sonnet 5, the minimum cacheable prompt is 1,024 tokens.

A large context ceiling still does not mean every request should fill the window. Relevance filtering, caching, retrieval, and explicit context management can reduce both cost and distraction.

Context capacity

Context window

1,000,000

Standard max output

128,000

Input + working stateOutput limit

Message Batches can expose up to 300K output tokens in beta. Sonnet 5 uses a 1,024-token minimum cacheable prompt.

05 / Thinking & effort

Adaptive Thinking Is On by Default—but Can Be Disabled

Claude Sonnet 5 made adaptive thinking the default Sonnet behavior while still allowing applications to turn thinking off completely when latency or workload simplicity makes reasoning unnecessary.

A Different Contract From Both Sonnet 4.6 and Sonnet 5.5

On Sonnet 4.6, a request that omitted thinking ran without thinking.

On Sonnet 5, the same request runs with adaptive thinking by default. The model decides when and how much to reason, while the effort parameter controls how much computational work it should invest.

Manual extended thinking is no longer accepted. Requests using thinking: {"type": "enabled", "budget_tokens": N} return a 400 error.

Unlike Sonnet 5.5, however, Sonnet 5 still accepts thinking: {"type": "disabled"}. That makes Sonnet 5 a distinctive compatibility point for applications that want Claude 5 capability but still need a fully non-thinking path for certain requests.

Anthropic supports five effort levels: low, medium, high, xhigh, and max. high is the default.

Adaptive thinking controls
disabled · supportedlowmediumhigh · defaultxhighmax

Latency-sensitive request

Use low effort or disable thinking when evaluations show that reasoning adds latency without improving the accepted result.

Hard coding or agentic task

Anthropic recommends xhigh for the most difficult coding and agentic workloads, with max reserved for cases where evals justify unrestricted effort.

06 / Tokenizer

Sonnet 5 Introduced a New Tokenizer

Claude Sonnet 5 changed tokenization enough that the same text produces approximately 30% more tokens than on Sonnet 4.6, depending on content.

Lower Token Prices Do Not Translate Directly Into the Same Percentage of Savings

The API request and response shapes did not change because of the tokenizer, but token-based assumptions did.

Anything that depends on token counts needs to be recalibrated:

  • prompt-size forecasts;
  • monthly cost estimates;
  • context-window thresholds;
  • compaction or trimming rules;
  • max_tokens values;
  • rate-limit planning;
  • benchmark comparisons.

The nominal context window remains 1M tokens, but if each token covers less text on average, the same 1M-token window can hold less equivalent source text than Sonnet 4.6.

Similarly, an output limit tuned closely around Sonnet 4.6 responses may truncate equivalent Sonnet 5 output.

Anthropic recommends recounting prompts with the token-counting API instead of carrying forward old estimates.

Token accounting changes
  1. 01

    Recount prompts

    Measure real production payloads under Sonnet 5 instead of reusing Sonnet 4.6 token counts.

  2. 02

    Revisit max_tokens

    Equivalent text can require more tokens, so output ceilings tuned for earlier Sonnet models may truncate responses.

  3. 03

    Recheck context thresholds

    Retrieval, trimming, and context-management logic based on old token ratios can trigger at the wrong point.

  4. 04

    Recalculate effective savings

    The $2/$10 rate is lower than Sonnet 4.6, but the tokenizer means equivalent-request savings are workload-dependent.

07 / Tools

A More Permissive Tool Contract Than Sonnet 5.5

Claude Sonnet 5 supports the same general tool and platform feature set as Sonnet 4.6, except Priority Tier, and it retains forced tool-choice patterns that Sonnet 5.5 later removes.

Useful for Existing Agents That Depend on Explicit Tool Forcing

Sonnet 5 accepts tool_choice modes that force any tool or a specific named tool.

That can matter in deterministic application flows where the product expects the model to call a tool rather than return free-form text.

Sonnet 5.5 changes this contract: forced any or named-tool choice returns a 400 error and applications must move toward automatic tool selection with strict schemas and stronger prompt instructions.

Sonnet 5 also predates several Sonnet 5.5-specific agent controls such as per-message effort, mid-conversation system messages, mid-conversation tool changes, and the lower 512-token prompt-cache minimum.

Assistant-message prefilling is not supported on Sonnet 5, as was already true for Sonnet 4.6.

Priority Tier is also unavailable on Sonnet 5.

Legacy Claude 5 tool contract
  • Tool use

    Supports server-side and client-defined tool workflows for agentic applications.

    Supported
  • Forced tool choice

    Can force any tool or a named tool, unlike Sonnet 5.5.

    Supported
  • Prompt caching

    Cache stable prompt prefixes with a 1,024-token minimum cacheable prompt.

    Supported
  • Image input

    Combine supported images with text and long-context workflows.

    Supported
  • Assistant prefill

    A prefilled final assistant turn returns an error.

    Not listed
  • Priority Tier

    Anthropic explicitly excludes Priority Tier from the Sonnet 5 platform feature set.

    Not listed

08 / Safeguards

The First Sonnet Model With Real-Time Cybersecurity Safeguards

Claude Sonnet 5 introduced real-time cybersecurity safeguards to the Sonnet tier, making refusal handling a production integration concern for security-sensitive workloads.

Refusal Is a Model Outcome, Not an HTTP Error

Anthropic documents that prohibited or high-risk cybersecurity requests can be declined by the model.

A refusal is returned as a successful HTTP 200 response with stop_reason: "refusal" rather than as a transport error.

That means production applications should inspect the response stop reason explicitly. A generic handler that treats every HTTP 200 as an ordinary completion may misclassify a safety outcome as usable model output.

This is particularly relevant for legitimate security products because benign defensive work can overlap with risky technical domains. Anthropic provides a Cyber Verification Program for eligible legitimate security work that needs reduced restrictions.

Sonnet 5.5 later expands safeguard categories and fallback behavior, but Sonnet 5 remains the first Sonnet baseline where real-time cyber refusal handling becomes part of the API integration story.

Real-time cyber safeguards
  1. 01

    Inspect stop_reason

    Treat refusal as a distinct model outcome even though the API request itself succeeds with HTTP 200.

  2. 02

    Track refusal rates

    Include policy refusals in production completion metrics because they affect accepted-result rates and fallback behavior.

  3. 03

    Separate policy from model failure

    Do not classify a safeguard refusal as a parsing, transport, or ordinary quality failure.

  4. 04

    Validate security workflows

    Legitimate cybersecurity products should evaluate safeguard behavior and provider verification options before production rollout.

09 / Sonnet 5 → 5.5

Migrating From Claude Sonnet 5 to Sonnet 5.5

Sonnet 5.5 keeps the same $2/$10 pricing and the same tokenizer as Sonnet 5, but changes thinking, forced tool use, conversation-state behavior, cache minimums, progress streaming, and several agent integrations.

The Economics Stay Stable While the Orchestration Contract Changes

The easiest part of the migration is token accounting: Sonnet 5.5 uses the same tokenizer and standard prices as Sonnet 5.

The harder part is API behavior.

Sonnet 5 can turn thinking fully off with thinking: disabled. Sonnet 5.5 rejects that setting and introduces between_tools as the lowest-thinking alternative.

Sonnet 5 accepts forced tool choice. Sonnet 5.5 rejects tool_choice: any and named-tool forcing.

The prompt-cache minimum also falls from 1,024 tokens on Sonnet 5 to 512 tokens on Sonnet 5.5.

Sonnet 5.5 adds per-message effort, mid-conversation system messages, and mid-conversation tool changes. It also binds preserved thinking blocks more tightly to model and conversation state, so append-only histories become safer.

Progress text changes shape too. On Sonnet 5, text written between tool calls is returned as ordinary text blocks. On Sonnet 5.5, longer progress notes can arrive as thinking blocks.

Computer-use and advisor-tool compatibility also change.

Migration differences
  1. 01

    Replace disabled thinking

    If the application disables thinking on Sonnet 5, migrate that path to between_tools or adaptive thinking on Sonnet 5.5.

  2. 02

    Remove forced tool selection

    Replace any-tool or named-tool forcing with automatic selection, strict tool schemas, and explicit prompting.

  3. 03

    Review conversation state

    Keep preserved thinking histories append-only and test progress streaming, computer use, and advisor integrations.

  4. 04

    Take advantage of newer controls

    Re-evaluate caching, per-message effort, dynamic tools, and mid-conversation instructions after the migration.

10 / Evaluation

Claude Sonnet 5 Strengths and Limitations

Claude Sonnet 5 remains a useful production and migration baseline because it combines current-scale context and low Sonnet pricing with a more permissive thinking and tool contract than Sonnet 5.5.

Strengths

  • Strong coding and agentic baseline

    Anthropic identifies coding and agentic tasks as the largest capability gains over Sonnet 4.6.

  • Low Sonnet-tier pricing

    $2/$10 standard pricing supports frequent production use without Opus-tier token costs.

  • Flexible thinking behavior

    Adaptive thinking runs by default but can still be completely disabled for latency-sensitive or simple workloads.

  • Permissive forced-tool support

    Existing agents can force any or named tools, a behavior removed in Sonnet 5.5.

What to consider

  • Legacy model

    Anthropic recommends migrating to Claude Sonnet 5.5 for improved performance.

  • New tokenizer complicates Sonnet 4.6 comparisons

    The same text uses roughly 30% more tokens than Sonnet 4.6, so historical token and cost assumptions need recalibration.

  • Fewer dynamic agent controls than Sonnet 5.5

    Per-message effort, mid-conversation system messages, dynamic tool changes, and a lower cache threshold arrive in the successor.

  • No Priority Tier

    Applications that depend on Priority Tier need a different model or deployment strategy.

Preserve the first Claude 5 Sonnet baseline

Compare Claude Sonnet 5 with Sonnet 5.5 on your own workload

Replay coding, analysis, tool-use, long-context, and latency-sensitive requests to measure how the newer Sonnet changes quality, thinking behavior, tool orchestration, caching, and total cost.

Start Free

Claude Sonnet 5 remains active as a legacy model, while Anthropic recommends migrating to Claude Sonnet 5.5 for improved performance.

Common Questions

What is Claude Sonnet 5?

Claude Sonnet 5 is Anthropic’s first Claude 5 Sonnet model, released on June 30, 2026 as a capability upgrade over Sonnet 4.6 with lower standard pricing and stronger coding and agentic performance.

How much does Claude Sonnet 5 cost?

Claude Sonnet 5 costs $2 per 1M input tokens and $10 per 1M output tokens. Five-minute cache writes cost $2.50/MTok, one-hour cache writes cost $4/MTok, and cache reads cost $0.20/MTok. Batch input and output receive a 50% discount.

What is the Claude Sonnet 5 context window?

Claude Sonnet 5 has a 1,000,000-token context window and supports up to 128,000 output tokens in standard requests.

Can Claude Sonnet 5 output more than 128K tokens?

Anthropic documents a beta maximum of up to 300,000 output tokens through the Message Batches API. Standard requests use the 128K output limit.

What is Claude Sonnet 5’s knowledge cutoff?

Anthropic lists January 2026 as both the reliable knowledge cutoff and the training-data cutoff.

Does Claude Sonnet 5 support images?

Yes. Claude Sonnet 5 accepts text and image input and produces text output.

Does Claude Sonnet 5 use adaptive thinking?

Yes. Adaptive thinking is on by default, with high as the default effort level. The model supports low, medium, high, xhigh, and max effort.

Can thinking be disabled on Claude Sonnet 5?

Yes. Claude Sonnet 5 accepts thinking type disabled. This differs from Claude Sonnet 5.5, where the disabled setting is rejected and between_tools is used to remove up-front thinking.

Does Claude Sonnet 5 support budget_tokens?

No. Manual extended thinking with thinking type enabled and budget_tokens was removed in Sonnet 5 and returns a 400 error. Use adaptive thinking with effort instead.

Why does Claude Sonnet 5 use more tokens than Sonnet 4.6?

Sonnet 5 introduced a new tokenizer. Anthropic states that the same text produces approximately 30% more tokens than on Sonnet 4.6, depending on content.

Can Claude Sonnet 5 force a tool call?

Yes. Sonnet 5 accepts forced any-tool and named-tool selection. Sonnet 5.5 removes that behavior, so forced-tool integrations require migration changes.

Does Claude Sonnet 5 support Priority Tier?

No. Anthropic states that Sonnet 5 supports the same tools and platform features as Sonnet 4.6 except Priority Tier.

What is the minimum cacheable prompt for Claude Sonnet 5?

Anthropic documents a 1,024-token minimum cacheable prompt for Sonnet 5. Sonnet 5.5 lowers that minimum to 512 tokens.

How is Claude Sonnet 5 different from Claude Sonnet 5.5?

Both use the same tokenizer, $2/$10 pricing, 1M context, and 128K standard output. Sonnet 5 allows thinking disabled and forced tool choice, while Sonnet 5.5 replaces disabled thinking with between_tools, removes forced tool selection, lowers the cache minimum, and adds newer dynamic agent controls.

Is Claude Sonnet 5 still available?

Yes. Anthropic lists Claude Sonnet 5 as Active (legacy). It was released June 30, 2026, with retirement not scheduled sooner than June 30, 2027.

Should I use Claude Sonnet 5 for a new application?

For new development, evaluate Claude Sonnet 5.5 first because Anthropic recommends the newer model. Sonnet 5 remains useful for existing integrations, forced-tool workflows, disabled-thinking compatibility, regression testing, and migration baselines.

Model information

Last updated

Specifications, pricing, lifecycle, thinking behavior, tokenizer changes, tool compatibility, cybersecurity safeguards, and migration details on this page are based on Anthropic’s official Claude Sonnet 5 overview, What’s New guide, prompting guidance, effort documentation, and Sonnet 5.5 migration guide.

Claude Sonnet 5 — Pricing, 1M Context, Adaptive Thinking & Agentic Coding | EidoStack