Anthropic announced Claude Haiku 5.5 on October 7, 2026, with adjustable effort and API prices that rise when a prompt exceeds 100,000 tokens. Standard rates start at $0.10 per million input tokens and $0.50 per million output tokens; above that threshold, they are $0.50 and $2.50. Anthropic’s announcement describes Haiku 5.5 as its first Haiku-class model with adjustable effort.
What changes above 100,000 tokens
Anthropic’s pricing documentation lists these standard first-party API rates in USD per million tokens. Tokens are the units the API bills for input and generated output.
| Charge | Prompt up to 100,000 tokens | Prompt over 100,000 tokens |
| Input | $0.10 per million tokens | $0.50 per million tokens |
| Output | $0.50 per million tokens | $2.50 per million tokens |
| Cache read | $0.01 per million tokens | $0.05 per million tokens |
| Cache write, five minutes | $0.125 per million tokens | $0.625 per million tokens |
| Cache write, one hour | $0.20 per million tokens | $1.00 per million tokens |
The listed rate is five times higher in the over-100,000-token band for every charge in the table. Anthropic also lists a 50% Batch API discount on input and output: $0.05 and $0.25 per million tokens in the lower band, or $0.25 and $1.25 in the higher band.
Global routing uses the standard first-party rates. Selecting US-only inference for Haiku 5.5 applies a 1.1× multiplier across token price categories.
What Anthropic’s average-savings estimate means
Anthropic estimates that Haiku 5.5 costs about 75% less to run on average than Haiku 4.5. Separately, it says Haiku 5.5’s token prices are 90% lower for prompts up to 100,000 tokens and 50% lower for prompts over that threshold. The 75% figure is an average workload estimate, while the other percentages describe token rates; none specifies the cost of a particular task. Anthropic also says its updated tokenizer uses slightly more tokens per task.
That distinction matters when you budget: the API’s bill depends on the tokens used, the applicable prompt-length band and any cache or endpoint choices—not just the input rate printed first.
Where Haiku 5.5 fits
Anthropic positions Haiku 5.5 for high-volume, cost-sensitive work such as summaries, classification, compaction, database queries and subagent tasks. Its adjustable effort setting lets developers change the effort level; Anthropic identifies Haiku 5.5 as the first model in the Haiku class to offer that control.
Anthropic’s model documentation lists a one-million-token context window and a standard maximum output of 128,000 tokens. The model’s ID is claude-haiku-5-5. Its context window and the 100,000-token pricing boundary are separate specifications.
Anthropic says Haiku 5.5 is available through Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. For related context on the model family, read our earlier coverage of Claude Sonnet 5.5.