Claude Haiku 5.5 is live at $0.10. The 100k cliff is the real price.
Anthropic shipped Haiku 5.5 on 7 October at $0.10 / $0.50 per million tokens under 100k. Over that line the rate is 5x. The new tokenizer spends about 30% more tokens.
Anthropic shipped Haiku 5.5 on 7 October at $0.10 / $0.50 per million tokens under 100k. Over that line the rate is 5x. The new tokenizer spends about 30% more tokens.
On 7 October 2026 Anthropic released Claude Haiku 5.5 on the API, Bedrock, Google Cloud, and Microsoft Foundry. Short prompts are 90% cheaper than Haiku 4.5. Average savings are closer to 75% after a tokenizer change, and prompts over 100,000 tokens jump to $0.50 / $2.50.
On 7 October 2026 Anthropic put Claude Haiku 5.5 on the Claude API. Model id is claude-haiku-5-5. It is also on Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Same day it landed in Claude Code and as a selectable model on claude.ai for Free, Pro, Max, Team, and Enterprise.
I read the launch as a routing decision, not a new chat personality. Short prompts are $0.10 per million input tokens and $0.50 per million output tokens. Haiku 4.5 was $1 and $5. Over 100,000 tokens the new rates jump to $0.50 and $2.50. Anthropic says about 90% of Haiku 4.5 requests sat under that line.
Haiku 5.5 is the small model in the Claude 5.5 family. Context window is 1 million tokens. Max output on the synchronous Messages API is 128,000 tokens. On the Message Batches API, beta header output-300k-2026-03-24 raises that to 300,000. Input is text and images. Output is text. Knowledge cutoff is June 2026. Retirement is not sooner than 7 October 2027.
It is the first Haiku with adaptive thinking and an effort parameter. Adaptive thinking is on by default. Default effort is medium. Omit temperature, top_p, and top_k. A non-default value returns a 400. Thinking blocks only work in the account that produced them, or an account linked to it.
Anthropic also cut Sonnet 5.5 cache reads in half, from $0.20 to $0.10 per million tokens. They say that knocks about 20% off Sonnet 5.5 on most agentic work, because cache reads are a large share of spend. This week Max and Team subscribers get a monthly Claude Platform credit: $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team. Credits work on any model. They are for trying API calls, not a standing production discount.
Python and TypeScript SDKs picked up beta support for computer use and browser use the same week. Anthropic points at Haiku 5.5 for those loops. At standard speed it is the fastest model in the current lineup. The announcement footnote: it is still slower than Opus models in Fast Mode.
Per million tokens:
Haiku 4.5 cache reads were $0.10. Cache writes were $1.25. No 100k split. Sonnet 5.5 stays at $2 input and $10 output, with the new $0.10 cache read.
The 90% cut is the headline. The invoice has two leaks.
First, the 100,000-token cliff. A prompt that tips over pays 5x on input and output. Stuff a repo, a long thread, or a retrieved document set into one call and you are on $0.50 / $2.50. That is still half of Haiku 4.5, not a tenth.
Second, the tokenizer. Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later. Anthropic says the same text counts as about 30% more tokens than on Haiku 4.5. Their average saving is about 75% once that inflation is in the math, not 90%. I would budget 75% until a week of my own logs says otherwise.
A classification call at 2,000 input tokens and 200 output tokens, short tier: $0.0002 input plus $0.0001 output, so $0.0003 before caching. At 100,000 calls a day that is about $30. The same shape on Haiku 4.5 was closer to $0.003 a call, about $300 a day, before the tokenizer difference. Cache the stable prefix and the input side drops to a tenth of the short-tier input rate.
Anthropic's own table. Their eval, not an outside lab.
Artificial Analysis, writing the same day, put Haiku 5.5 at max effort slightly ahead of GPT-6 Luna on its Intelligence Index, and noted the model burns far more output tokens at max effort: about 162k output tokens per index task versus about 50k for Luna. Their Terminal-Bench 4.0 read was 33%, not Anthropic's 39.2%. The jump from Haiku 4.5 looks real. Not a reason to skip your own set.
Haiku 5.5 is near Sonnet on some narrow scores and well behind on hard terminal work. Anthropic says Sonnet 5.5 and Opus 5.5 stay the pick for complex agentic coding. Haiku 5.5 is for compaction, summarization, classification, routing, and subagent chores that were too expensive to run on every turn.
Vendor-selected quotes from the launch post, not a public benchmark.
Asana ran it through an eval suite for AI Teammates. Aaron Vinh said they saw over a 30% cut in latency for task completions and up to 2.5x faster inference per agent turn versus the model they use today.
HubSpot tested smaller models on simulated CRM portals. Ze'ev Klapow said Haiku 5.5 scored 92.8% averaged over three runs, the best on that suite, and won a stale-record audit on speed, hit rate, and false positives.
AlphaSense ran 400 queries on Ask in Document, about 8 million calls a week. Daniel Campos reported 0.84 versus 0.76 on Haiku 4.5.
Box said early testing scored 11 points higher than Haiku 4.5 at about half the latency. Cognition put Haiku 5.5 in as a sidekick in Devin Fusion, Opus 5.5 as the lead, and said Fusion held a FrontierCode score of 66.2 at lower cost and latency.
A hint. Not a reason to swap a production router without a shadow week.
Cyber safeguards are tighter than Haiku 4.5 and looser than Sonnet 5.5. Anthropic says they allow a wider set of defensive tasks than Sonnet 5.5, and still block penetration testing and attacker-style technique. Biology safeguards match Sonnet 5, Sonnet 5.5, and Opus 5. Wider work goes through the verification programs. Reuters noted the narrow cyber block and said most everyday tasks are unaffected.
Support, extraction, and routing will not feel this. An offensive-security agent will refuse part of the job.
Context is 1 million tokens, up from 200k on Haiku 4.5. Long prompts hit the higher price band, and the new tokenizer spends the window faster. Docs put 1 million tokens at roughly 555,000 words on the current tokenizer.
Batch is the boring win. Fifty percent off, plus the 300k output beta, fits overnight compaction and labeling. Do not put a live support turn on batch.
Keep Sonnet or Opus on the turn that has to plan, edit a large diff, or recover from a failed tool call. Send Haiku 5.5 the turns that repeat: classify the ticket, pull one field, summarize the thread, compact history, fan out a subagent that reads one doc.
Pin claude-haiku-5-5. Log input tokens, output tokens, cache reads, and whether the request crossed 100,000 tokens. If more than a small slice crosses that line, the 90% story is not your story.
Set effort to low on classification and extraction. Leave medium only where you have already seen misses. Max effort is where outside testers saw the token bill jump.
Shadow it against Haiku 4.5 for a few days on the real queue. Compare latency, refusal rate, and dollars per successful task. The migration guide is on the Claude Platform docs. Pricing page and system card are the other two tabs worth open before you flip the default.
Sources: Anthropic launch post, 7 October 2026, https://www.anthropic.com/claude-haiku-5-5. Model overview, https://platform.claude.com/docs/en/models/haiku-5-5/overview.
Community
0 comments
React to this article
Trending now
Written by
Kirtesh Admute
Founder
Kirtesh Admute is the founder of IndieFounder, a platform for founders, builders, and people curious about technology. He writes about AI, startups, software, product building, and the lessons that come from building in public.
See an issue with this story?