Anthropic efficiency model
Claude Haiku 4.5
Anthropic’s fastest current model for interactive applications, high-volume processing, cost-sensitive automation, and lightweight agent workflows that still need strong reasoning, vision, and tool use.
- Context window
- 200K
- tokens
- Max output
- 64K
- tokens
- Input
- $1.00
- per 1M tokens
- Cache read
- $0.10
- per 1M tokens
- Output
- $5.00
- per 1M tokens
01 / Overview
What Claude Haiku 4.5 Is
Claude Haiku 4.5 is Anthropic’s fastest current model, designed for real-time applications, high-volume processing, and cost-sensitive deployments that still need strong reasoning, coding, vision, and tool-use capability.
Near-Frontier Capability at the Lowest Current Claude Price
Anthropic released Claude Haiku 4.5 on October 15, 2025 and describes it as the fastest and most intelligent Haiku model with near-frontier performance.
Its role in the Claude lineup is different from Sonnet and Opus.
Haiku is not the model you choose because a task is maximally difficult. It is the model you evaluate first when latency, throughput, and unit economics are central constraints.
The model provides a 200,000-token context window and supports up to 64,000 output tokens. It accepts text and images and produces text.
Haiku 4.5 also supports manual extended thinking for workloads that need more reasoning. Unlike Sonnet 4.6 and newer Claude generations, it does not support adaptive thinking or the effort parameter.
Anthropic currently lists Haiku 4.5 as Active (latest) and classifies its comparative latency as Fastest.
- Model ID:
claude-haiku-4-5-20251001. - Alias:
claude-haiku-4-5. - Released October 15, 2025.
- Status: Active (latest).
- Retirement not sooner than October 15, 2026.
- 200K context window.
- 64K maximum output.
- Text and image input with text output.
- Manual extended thinking supported.
- Provider
- Anthropic
- Family
- Claude Haiku 4.5
- Model ID
- claude-haiku-4-5-20251001
- Reliable knowledge cutoff
- Feb 2025
- Training cutoff
- Jul 2025
- Comparative latency
- Fastest
- Lifecycle
- Active · latest
02 / High-volume use
High-Volume Workloads Where Latency and Cost Compound
Haiku 4.5 is most valuable when an application makes enough calls that per-request latency and token cost materially affect user experience or gross margin.
Optimize the Cost of a Useful Result
Low token price is not enough by itself.
A cheap model that requires frequent retries, produces more invalid outputs, or needs a stronger fallback on most requests can become more expensive than a larger model.
The correct evaluation metric is therefore cost per accepted result.
For high-volume workloads, measure success rate, retry rate, response latency, output length, tool-call accuracy, fallback frequency, and total tokens consumed.
Typical Haiku candidates include classification, extraction, moderation support, routing, summarization, lightweight coding assistance, customer-facing chat, background automation, and repetitive tool-based operations.
The same logic applies to interactive UX. If a model responds faster but misses important instructions often enough to trigger correction loops, the latency advantage may disappear.
- 01
Classification and routing
Categorize requests, select downstream workflows, score content, or route difficult cases to a stronger model.
- 02
Extraction and transformation
Convert text or images into structured application data at high request volume.
- 03
Interactive product features
Use the fastest Claude tier where perceived latency directly affects user experience.
- 04
Background automation
Process queues, summaries, metadata, and repetitive AI operations where throughput and unit cost matter.
03 / Model routing
Use Haiku as an Efficiency Tier, Not a Universal Model
A practical production architecture can route routine requests to Haiku 4.5 and escalate only the difficult minority to Sonnet, Opus, or Fable.
Escalation Can Be Cheaper Than Premium-Only Routing
Suppose most requests are straightforward but a small fraction require deep reasoning, long-horizon coding, or a million-token working context.
Routing every request to the strongest model pays premium-model prices even when that capability is unnecessary.
A Haiku-first architecture can instead:
- attempt the task on Haiku;
- validate the result;
- escalate only if quality, confidence, or structural checks fail.
The savings depend on the fallback rate. If Haiku succeeds on 90% of requests and expensive escalation is rare, routing can materially lower average inference cost.
If Haiku fails often and every request requires a second model call, the architecture may cost more and add latency.
That is why routing policy should be evaluation-driven. Use representative production data and track the combined economics of the full decision tree.
- 01
Try Haiku
Run the request on the fastest, lowest-cost current Claude tier.
- 02
Validate
Check structured output, task-specific quality criteria, confidence signals, or deterministic business rules.
- 03
Escalate selectively
Route only failed or unusually difficult requests to Sonnet, Opus, or Fable.
- 04
Measure blended cost
Evaluate the complete routing tree, including retries and premium-model fallbacks, rather than Haiku token price alone.
04 / Pricing
Claude Haiku 4.5 Pricing
Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens, making it the lowest-priced current model in Anthropic’s main Claude lineup.
Low Base Cost Plus Prompt Caching and Batch Discounts
Anthropic lists 5-minute cache writes at $1.25 per million tokens, 1-hour cache writes at $2 per million, and cache reads at $0.10 per million.
Batch API requests receive a 50% discount on standard input and output pricing, reducing the rate to an effective $0.50 input and $2.50 output per million tokens for eligible batch usage.
For high-volume systems, small price differences can compound quickly.
At the same time, manual extended thinking is billed as output usage. A reasoning-heavy Haiku request can therefore cost more than a short ordinary request even though the base model remains inexpensive.
Measure the full request shape: prompt tokens, cache writes, cache reads, thinking output, visible output, retries, and fallback calls.
- Input: $1.00 per 1M tokens.
- Output: $5.00 per 1M tokens.
- 5-minute cache write: $1.25 per 1M tokens.
- 1-hour cache write: $2.00 per 1M tokens.
- Cache read: $0.10 per 1M tokens.
- Batch API: 50% discount on input and output.
1M tokens · USD
- Input
- $1.00
- 5m cache write
- $1.25
- 1h cache write
- $2.00
- Cache read
- $0.10
- Output
- $5.00
Example: 20K input + 4K output
- Input cost
- $0.0200
- Output cost
- $0.0200
- Estimated total before extra thinking output
- $0.0400
05 / Context
200K Context and 64K Output
Claude Haiku 4.5 provides a 200,000-token context window and supports up to 64,000 output tokens.
Enough Capacity for Many Production Tasks, but Not a 1M-Context Model
A 200K window can hold substantial documents, conversation history, code, tool results, images, system instructions, and retrieved context.
For many classification, support, extraction, summarization, and lightweight agent workloads, that is more capacity than the request actually needs.
However, current Sonnet, Opus, and Fable tiers provide 1M-token context windows and 128K standard output.
That difference matters for very large repositories, long document collections, persistent agents, or research tasks that accumulate substantial working state.
Haiku therefore works best when the application can keep context focused. Retrieval, selective history, and external memory can preserve the latency and cost advantages of the smaller model instead of filling every request with unnecessary data.
Context window
200,000
Max output
64,000
Haiku 4.5 has a smaller context and output envelope than current Sonnet, Opus, and Fable models, which use 1M context and 128K standard output.
06 / Extended thinking
Manual Extended Thinking Without Adaptive Effort
Claude Haiku 4.5 supports manual extended thinking with budget_tokens, but it does not support adaptive thinking or the effort parameter.
Turn Reasoning On Only Where It Earns Its Cost
Extended thinking is enabled with:
thinking: {"type": "enabled", "budget_tokens": N}
This gives the application direct control over the reasoning-token budget.
Anthropic specifically recommends considering extended thinking for coding and reasoning tasks where the additional inference can produce significant performance improvements.
For routine classification, extraction, routing, or straightforward transformations, reasoning may add unnecessary latency and output cost.
That makes Haiku especially suitable for selective thinking: keep ordinary high-volume requests fast, and enable a measured thinking budget only for task classes where evaluation shows a real quality gain.
Haiku 4.5 does not support adaptive thinking. It also does not support effort, so there is no low, medium, high, or max effort ladder.
High-volume routine task
Keep thinking off when the workload passes quality thresholds without extra reasoning.
Coding or difficult reasoning
Enable extended thinking and tune budget_tokens when evaluation shows a meaningful improvement.
07 / Tools
Tool Use With a Simpler Thinking Model
Claude Haiku 4.5 supports tool use, but its manual-thinking behavior differs from larger Claude 4.5 and newer models in one important way: interleaved thinking is not supported.
Haiku Can Chain Tools Without Reasoning Between Tool Results
When extended thinking is enabled, Claude can reason before selecting a tool and can use tools as part of the request.
As with other manual extended-thinking models, forced tool choice is restricted while thinking is active. tool_choice: auto and tool_choice: none are supported, while forced any or named-tool selection is incompatible with extended thinking.
During tool-use continuations, the application must return the complete unmodified thinking blocks from the most recent assistant message to preserve reasoning continuity.
However, Haiku 4.5 does not support interleaved thinking. Adding the interleaved-thinking-2025-05-14 beta header is accepted on the Claude API but ignored for this model.
That does not prevent multiple tool calls. Haiku can still chain tools. It simply does not insert the manual reasoning pattern between tool results that supported interleaved-thinking models can use.
- Supported
Tool use
Use client-defined and supported server tools in application workflows.
- Supported
Extended thinking with tools
Reason before tool selection and preserve the latest thinking blocks when continuing tool turns.
- Not listed
Interleaved thinking
Haiku 4.5 does not reason in manual thinking blocks between tool calls; the legacy beta header is ignored.
- Not listed
Forced tools during thinking
When thinking is enabled, use automatic or no-tool selection rather than forced any-tool or named-tool choice.
- Supported
Image input
Combine images with text for visual extraction, analysis, and multimodal automation.
- Supported
Streaming
Stream text and supported thinking content for interactive applications.
08 / Caching
Prompt Caching Has a 4,096-Token Minimum
Claude Haiku 4.5 supports prompt caching, but its minimum cacheable prompt is 4,096 tokens—larger than the 1,024-token threshold used by several Sonnet models.
Caching Is Most Useful for Substantial Repeated Prefixes
Short prompts marked with cache_control are not cached if they do not meet the minimum length. Anthropic does not return an error; the request simply processes normally without creating a cache entry.
That means Haiku caching is most relevant for larger repeated prefixes: substantial system instructions, long reference documents, tool definitions, policies, or other stable context.
Haiku also has distinct thinking-block retention semantics.
For all Haiku models, prior thinking blocks are removed from context calculations except where they must be preserved for the latest tool-use continuation. This differs from newer Opus and Sonnet generations that preserve more historical thinking state by default.
Changing thinking settings or budget allocation can also invalidate message cache breakpoints, although system prompts and tools can remain cached.
These details matter in high-volume systems because a cache architecture that works well with Sonnet may not produce identical hit behavior with Haiku.
- 01
4,096-token minimum
Shorter prompts cannot be cached even when cache_control is present.
- 02
$0.10 cache reads
Large stable prefixes become inexpensive to reuse once a valid Haiku cache entry exists.
- 03
Last-turn thinking continuity
Haiku uses older thinking-block retention semantics rather than preserving all prior thinking turns by default.
- 04
Thinking changes can break message caches
Changing enabled/disabled thinking or budget allocation can invalidate message cache breakpoints.
09 / Migration
Migrating From Claude Haiku 3.5 to Haiku 4.5
Haiku 4.5 is the current Haiku generation and upgrades the efficiency tier with stronger intelligence, 64K output, manual extended thinking, and improved speed while keeping the model focused on low-latency, high-volume deployment.
Re-Evaluate Rather Than Only Swap the Model ID
The Claude API model changes from claude-3-5-haiku-20241022 to claude-haiku-4-5-20251001.
Anthropic instructs developers to review new rate limits because Haiku 4.5 has separate limits from Haiku 3.5.
Haiku 4.5 also expands maximum output to 64K tokens and supports manual extended thinking for coding and reasoning workloads.
Prompt-caching assumptions need review as well. Haiku 4.5 has a 4,096-token minimum cacheable prompt, while Anthropic documents a smaller threshold for Haiku 3.5.
For applications that historically used Haiku only for simple classification, the most important migration opportunity may be capability rather than compatibility: some tasks previously routed to a larger model may now pass on Haiku 4.5.
Re-run the routing eval set after migration instead of preserving old model-tier assumptions.
- 01
Update the model ID
Move from claude-3-5-haiku-20241022 to the pinned claude-haiku-4-5-20251001 snapshot.
- 02
Review rate limits
Haiku 4.5 uses separate provider rate limits, so update throughput assumptions before production rollout.
- 03
Retest reasoning workloads
Use manual extended thinking on coding and reasoning cases that previously required a larger model.
- 04
Rebuild the routing threshold
Check whether improved Haiku capability can absorb workloads that were formerly escalated to Sonnet.
10 / Evaluation
Claude Haiku 4.5 Strengths and Limitations
Haiku 4.5 is an efficiency specialist: its strongest production case is achieving sufficient quality at lower latency and cost, not matching the hardest-task ceiling of Sonnet, Opus, or Fable.
Strengths
Fastest Claude latency
Anthropic classifies Haiku 4.5 as the fastest current model, making it a natural candidate for interactive and latency-sensitive features.
Lowest current Claude pricing
$1 input and $5 output per million tokens support high-volume workloads and low-cost routing tiers.
Near-frontier intelligence
Haiku 4.5 is designed to handle substantially more than simple classification while retaining efficiency-first economics.
Optional deeper reasoning
Manual extended thinking can improve difficult coding or reasoning tasks without forcing every routine request to pay reasoning latency.
What to consider
Smaller context than current premium tiers
200K context and 64K output are below the 1M/128K profile of current Sonnet, Opus, and Fable models.
No adaptive thinking or effort
Reasoning must be manually enabled and budgeted; the newer adaptive-thinking controls are unavailable.
No interleaved thinking
Haiku can use and chain tools but does not support manual reasoning blocks between tool results.
Large cache minimum
The 4,096-token minimum means short recurring prompts do not benefit from prompt caching.
Find the cheapest model that passes
Test Claude Haiku 4.5 on your production workload
Compare latency, accepted-result quality, token usage, cache behavior, and total task cost across extraction, classification, coding assistance, support, tool calls, and routing workloads before escalating to Sonnet or Opus.
Start FreeHaiku is most valuable when its speed and price advantages survive real workload evaluation—not when it is selected only because it is the cheapest Claude tier.
Common Questions
What is Claude Haiku 4.5?
Claude Haiku 4.5 is Anthropic’s fastest current model, designed for real-time applications, high-volume processing, and cost-sensitive deployments while retaining strong reasoning, vision, and tool-use capability.
What is the Claude Haiku 4.5 model ID?
The pinned Claude API model ID is claude-haiku-4-5-20251001. Anthropic also provides the convenience alias claude-haiku-4-5, which resolves to that pinned snapshot.
How much does Claude Haiku 4.5 cost?
Claude Haiku 4.5 costs $1 per 1M input tokens and $5 per 1M output tokens. Five-minute cache writes cost $1.25/MTok, one-hour cache writes cost $2/MTok, and cache reads cost $0.10/MTok. Batch input and output receive a 50% discount.
What is the Claude Haiku 4.5 context window?
Claude Haiku 4.5 has a 200,000-token context window and supports up to 64,000 output tokens.
What is Claude Haiku 4.5’s knowledge cutoff?
Anthropic lists February 2025 as the reliable knowledge cutoff and July 2025 as the training-data cutoff.
When was Claude Haiku 4.5 released?
Anthropic released Claude Haiku 4.5 on October 15, 2025. It is currently listed as Active (latest), with retirement not sooner than October 15, 2026.
Does Claude Haiku 4.5 support images?
Yes. Claude Haiku 4.5 accepts text and image input and produces text output.
Does Claude Haiku 4.5 support extended thinking?
Yes. Haiku 4.5 supports manual extended thinking with thinking type enabled and budget_tokens. Anthropic specifically recommends considering it for coding and reasoning workloads where extra reasoning improves performance.
Does Claude Haiku 4.5 support adaptive thinking or effort?
No. Haiku 4.5 uses manual extended thinking only. Adaptive thinking and the effort parameter are not supported.
Does Claude Haiku 4.5 support interleaved thinking?
No. Anthropic states that Haiku 4.5 does not support interleaved thinking. The interleaved-thinking beta header may be accepted by the Claude API but is ignored for this model.
Can Claude Haiku 4.5 use tools while extended thinking is enabled?
Yes. Tool use can be combined with extended thinking, but forced any-tool or named-tool choice is incompatible with thinking. Use automatic or no-tool selection and return the complete latest thinking blocks during tool-result continuation.
What is the minimum cacheable prompt for Claude Haiku 4.5?
Anthropic documents a 4,096-token minimum cacheable prompt for Haiku 4.5. Shorter prompts are processed normally but are not cached.
How is Claude Haiku 4.5 different from Claude Sonnet 4.5?
Both use manual extended thinking and have 200K context with 64K output, but Haiku 4.5 is optimized for the fastest latency and costs $1/$5 compared with Sonnet 4.5 at $3/$15. Haiku also does not support interleaved thinking.
How is Claude Haiku 4.5 different from Claude Sonnet 5.5?
Haiku 4.5 is the lower-cost, fastest tier with 200K context, 64K output, and manual extended thinking. Sonnet 5.5 provides 1M context, 128K output, adaptive thinking, newer agent controls, and stronger capability at $2/$10.
Should I choose Claude Haiku 4.5 or a larger Claude model?
Start with workload requirements. Haiku 4.5 is a strong candidate when latency, throughput, and cost matter and it meets your quality threshold. Escalate to Sonnet, Opus, or Fable when difficult reasoning, long-horizon work, or larger context consistently exceeds Haiku’s capabilities.
Model information
Last updated
Specifications, pricing, lifecycle, extended-thinking behavior, tool-use constraints, prompt-caching limits, latency positioning, and migration guidance on this page are based on Anthropic’s official Claude Haiku 4.5 overview, pricing documentation, extended-thinking documentation, prompt-caching documentation, release notes, and Haiku 4.5 migration guide.