Anthropic releases Claude Haiku 5.5 and cuts Sonnet cache-read prices
Anthropic’s new Haiku 5.5 lowers list rates for short prompts, while Sonnet 5.5 cache reads now cost half as much.
By Wei-Lin Zhao · AI Correspondent
· 3 min read
Anthropic released Claude Haiku 5.5 on Oct. 7 with lower list prices for its small-model tier, while reducing Claude Sonnet 5.5 cache-read pricing by 50%, according to reports by The Decoder. For buyers comparing Claude Haiku 5.5 pricing, the model costs $0.10 per million input tokens and $0.50 per million output tokens for prompts of up to 100,000 tokens.
That tier is 90% below the reported Haiku 4.5 rates of $1 per million input tokens and $5 per million output tokens. Above 100,000 tokens, Haiku 5.5 is priced at $0.50 per million input tokens and $2.50 per million output tokens, a 50% reduction from the prior model’s listed rates.
The launch fills out Anthropic’s 5.5 model family after the company introduced Opus 5.5 on Sept. 22. The company had said then that Sonnet 5.5 and Haiku 5.5 would follow in subsequent weeks.
What does Claude Haiku 5.5 cost?
- Prompts of up to 100,000 tokens: $0.10 per million input tokens and $0.50 per million output tokens.
- Prompts above 100,000 tokens: $0.50 per million input tokens and $2.50 per million output tokens.
- Cache reads: $0.01 per million tokens at the lower prompt tier and $0.05 at the higher tier, The Decoder reported.
- Cache writes: $0.125 per million tokens at the lower tier and $0.625 at the higher tier.
For a workload with one million input tokens and one million output tokens at the lower tier, those list rates add to $0.60, compared with $6 under the reported Haiku 4.5 input and output rates. Actual application bills will depend on the input-output mix, prompt lengths and use of cached context. The Decoder also reported that Haiku 5.5 uses an updated tokenizer that can consume somewhat more tokens per task, which can narrow savings relative to a per-token comparison.
Anthropic said the model is intended for high-volume work such as summarization, classification, data queries, customer support and narrowly defined subagent tasks. Reports said it is available through Anthropic’s API and via AWS, Google Cloud and Microsoft Azure.
The company reported large gains over Haiku 4.5 on its own tests. On the OSWorld 2.1 offline subset, Haiku 5.5 scored 72.4%, versus 15.7% for its predecessor, according to The Decoder. Anthropic also reported a 39.2% score on Terminal-Bench 4.0, compared with 0.0% for Haiku 4.5. Those are company-reported results, rather than independent evaluations. The cited comparisons also show Haiku 5.5 behind Sonnet 5.5, and Anthropic positions Sonnet and Opus for more complex agentic coding.
How does the Sonnet 5.5 cache-read cut affect costs?
Sonnet 5.5 cache reads now cost $0.10 per million tokens, down from $0.20, The Decoder reported. Cache reads apply when an application reuses prompt context that has already been stored, so the price cut does not halve the bill for every Sonnet workload. Anthropic said the change should make most agentic tasks about 20% cheaper, a figure that depends on how much of a workload consists of cached context.
This story draws on original reporting from SiliconANGLE.