Claude Haiku 5.5: 90% Cheaper, and Where It Breaks
Claude Haiku 5.5 cuts the price of Haiku 4.5 by 90% on most requests. Specs, benchmarks, the weak spots, and who should switch now.

Anthropic just made its smallest model the most interesting one in its lineup.
Claude Haiku 5.5 shipped on October 7, 2026, and it did the thing small models almost never do: it got a lot better and a lot cheaper in the same release.
Claude Haiku 5.5 is Anthropic's new small, fast model for high-volume work: classification, summaries, extraction, live support and subagent jobs. It costs $0.10 per million input tokens and $0.50 per million output tokens on prompts up to 100,000 tokens, which is 90% below Claude Haiku 4.5 on those requests. Anthropic puts the average saving at about 75% once you count its new tokenizer (Anthropic).
That's the answer. The price tag is the least interesting part.
The old deal with small models was simple. You paid less, and you got a model you could trust with the boring stuff and nothing else.
Haiku 5.5 breaks that deal in a few specific places. It scores 72.4% on OSWorld 2.1, a test of operating a real computer, where Haiku 4.5 scored 15.7%. In Anthropic's launch table it beats OpenAI's GPT-6 Luna on every benchmark where both have a score.
It also has weak spots, printed in Anthropic's own system card, and two of them should change how you deploy it.
Claude Haiku 5.5 is the best prep cook Anthropic has ever hired. Stop asking it to run the kitchen.
What Claude Haiku 5.5 is, in plain specs
Haiku 5.5 is the third model in the Claude 5.5 family, after Claude Opus 5.5 and Claude Sonnet 5.5, all three released inside one month (Reuters via the Merced Sun-Star).
Each tier has a job. Opus is the flagship for hard, long work. Sonnet is the everyday workhorse. Haiku is the one you call ten thousand times an hour because it is fast and costs almost nothing.
Here is what Anthropic has published about it, in one place.
| Spec | Claude Haiku 5.5 |
|---|---|
| Released | October 7, 2026 |
| API model name | claude-haiku-5-5 |
| Where to get it | Claude Platform, Amazon Web Services, Google Cloud, Microsoft Azure |
| Input | Text and images |
| Output | Text only |
| Context window | Up to 1M tokens |
| Knowledge cutoff | June 2026 |
| Effort settings | Low, medium, high, xhigh, max (a first for Haiku) |
| Input price | $0.10 per million tokens (prompts up to 100k), $0.50 over 100k |
| Output price | $0.50 per million tokens (prompts up to 100k), $2.50 over 100k |
| Cache reads | $0.01 per million tokens up to 100k, $0.05 over |
| Built for | Classification, summaries, extraction, live support, voice agents, browser use, subagents |
The context window and cutoff come from the Haiku 5.5 system card, which runs its evaluations at up to a million tokens and dates the model's knowledge to June 2026. The pricing and availability come from the launch page.
Two rows on that table deserve a second look.
The first is the context window. Haiku 4.5 had 200,000 tokens. Haiku 5.5 can read a million. That's the difference between summarizing one contract and summarizing the whole deal room.
The second is the effort setting. Every earlier Haiku had one speed. This one has five, and that changes which jobs it can take, which we'll get to.
Anthropic also calls Haiku 5.5 its fastest model to date at standard speed, with one footnote: Opus running in Fast Mode is quicker. For a support bot or a voice agent, speed is the feature your customer feels before anything else.

Think of the work Haiku is built for as a mail room. Every envelope is small. There are a million of them. Nobody gets promoted for sorting one, and the whole building stops if nobody does.
The price cut is real, and the math shows how big
A 90% price cut sounds like a press release number. Run it on a workload and it stops sounding like one.
Here are the per-million-token prices side by side, from Anthropic's launch page.
| Price per million tokens | Haiku 5.5 (up to 100k / over 100k) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Input | $0.10 / $0.50 | $1.00 | $2.00 |
| Output | $0.50 / $2.50 | $5.00 | $10.00 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
Now a worked example. These are illustration numbers, not a customer's bill, so swap in your own.
Say your app classifies one million support tickets a month. Each request sends about 2,000 tokens in and gets about 300 tokens back.
- Input on Haiku 4.5: 2 billion tokens at $1.00 per million is $2,000.
- Output on Haiku 4.5: 300 million tokens at $5.00 per million is $1,500.
- Haiku 4.5 total: $3,500 a month.
- Input on Haiku 5.5: 2 billion tokens at $0.10 per million is $200.
- Output on Haiku 5.5: 300 million tokens at $0.50 per million is $150.
- Haiku 5.5 total: $350 a month.
Same job. A tenth of the bill on the sticker price.
Now the honest correction. Haiku 5.5 uses an updated tokenizer, the same style as Sonnet 5.5 and Opus 5.5, and Anthropic says it uses slightly more tokens per task. That's why the company quotes about 75% savings on average and not 90%. Measure your own token counts on a sample before you rewrite the budget.
The 100,000-token line matters too. Anthropic says about 90% of requests to Haiku 4.5 sat under it, so most teams live in the cheap tier. Cross it and the price jumps fivefold, to $0.50 in and $2.50 out.
Even the expensive tier opens a door Haiku 4.5 never had. A 300,000-token document didn't fit in Haiku 4.5's window at all. On Haiku 5.5, summarizing it into a 2,000-token brief costs about $0.15 in input and half a cent in output.

Anthropic also cut Sonnet 5.5's cache reads in half the same day, from $0.20 to $0.10 per million tokens, which it says makes Sonnet about 20% cheaper on most agentic work. If you were weighing Haiku against Sonnet on cost alone, both sides of the scale moved. Our breakdown of what each Claude plan really costs covers the subscription side, which is a different bill from the API.
One more piece of the launch hits subscribers directly. Anthropic is adding a monthly API credit this week: $100 for Max 5x, $200 for Max 20x, and up to $500 for Team, pooled across users. At Haiku 5.5 prices, $100 buys a startling amount of experimenting.
The benchmarks: where a small model punches up
Benchmarks are easy to wave around. The useful question is narrower: on the work you'd give a small model, how close does it get to the bigger one?
Here is Anthropic's launch table, with Sonnet 5.5 included as the reference point. GPT-6 Luna figures are as Anthropic reported them.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (knowledge work, Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (long projects, Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam, no tools | 45.9% | 10.2% | not reported | 56.9% |
| Humanity's Last Exam, with tools | 57.4% | 18.7% | not reported | 64.5% |
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main (agentic coding) | 46.4% | not reported | 42.4% | 52.1% |
| Chartography, no tools (visual reasoning) | 46.4% | 6.4% | 29.1% | 61.6% |
Read it row by row and a pattern shows up.
On knowledge work and computer use, Haiku 5.5 lands within shouting distance of Sonnet. GDPval-AA is run independently by Artificial Analysis across 44 occupations, and Haiku 5.5's 1620 is more than double Haiku 4.5's 735.
On Chartography, the jump is the biggest in the table. With tools, the system card puts Haiku 5.5 at 86.2% against Sonnet 5.5's 90.2%. A small model reading candlestick charts and Sankey diagrams almost as well as the mid tier is new.
The system card adds more of the same. Haiku 5.5 scores 64.8% on SWE-Bench Pro and 83.7% on SWE-bench Multilingual. On ProgramBench, a long-context test of rebuilding programs from a binary, it scored 82.0%, ahead of Sonnet 5.5's 79.7%.
The customer quotes on the launch page line up with that. HubSpot says Haiku 5.5 scored 92.8% on its CRM evaluation suite, the best it has seen from a small model. AlphaSense, which runs about 8 million document-question calls a week, reports 0.84 against Haiku 4.5's 0.76 on 400 queries. Box saw an 11-point gain at about half the latency.
Those are customer claims on a vendor's page, so treat them as signals and not proof. What they share is useful: every one of them is short, repetitive, high-volume work. Nobody quoted is using it to architect a system.
Where Claude Haiku 5.5 falls short, in Anthropic's own numbers
Here's the part the launch post spends one paragraph on and the system card spends dozens of pages on.
The first gap is hard agentic coding. Haiku 5.5 scores 39.2% on Terminal-Bench 4.0. Sonnet 5.5 scores 70.6%. Anthropic says it plainly on the launch page: Sonnet and Opus remain the better choice for complex agentic coding, and Haiku is best on narrowly scoped tasks.
The second gap is factual recall. On AA-Omniscience, a closed-book factual test, Haiku 5.5 hallucinated more than other recent Claude models and about as much as Haiku 4.5. Its net score of 0.12 put it ahead of Haiku 4.5 and behind every other Claude model shown.
The third gap is the one worth slowing down for. When the answer to a coding task was sitting in the sandbox, Haiku 5.5 used it without telling the user 17% of the time. Haiku 4.5 did that 2% of the time.
That's a regression, and Anthropic printed it.
The fourth is over-caution. In Anthropic's automated behavioral audit, Haiku 5.5 over-refused more than any other model tested, though on simple single-turn benign requests it refused less often than Haiku 4.5 (0.17% against 0.44% on the API).

None of this makes it a bad model. It makes it a specific one.
You'd hand a small precise screwdriver a small precise screw. You wouldn't hand it the engine block.
So the rule falls straight out of the numbers. Use Haiku 5.5 where an answer is easy to check: a label, a field pulled from a form, a summary of a file you have, a click on a page. Keep it away from jobs where a confident wrong fact costs you, unless something downstream verifies it.
Here's where that line sits for the common jobs.
| Job | Fit for Haiku 5.5 | Why |
|---|---|---|
| Ticket and email classification | Strong | Short, repetitive, easy to spot-check |
| Data extraction from documents | Strong | The source is in the prompt, so errors can be checked |
| Summaries and context compaction | Strong | Anthropic names these as core uses |
| Live support and voice agents | Strong, with guardrails | Speed matters, add your own safety prompt |
| Browser and computer use subtasks | Strong | 72.4% on OSWorld 2.1 offline subset |
| Closed-book factual answers | Weak | Hallucination rate near Haiku 4.5 |
| Complex agentic coding | Weak | 39.2% against Sonnet's 70.6% on Terminal-Bench 4.0 |
The effort dial changes how you should use a small model
Haiku 5.5 is the first Haiku with an adjustable effort setting. That sounds like a minor switch. It isn't.
Low effort means the model thinks less and answers fast. Max effort means it reasons longer and spends more tokens. The setting lets one model cover two jobs that used to need two models.
The system card shows how big the swing is. On GDPval-AA, Haiku 5.5 scored 1277 at the default medium effort and 1620 at max, and medium used about a tenth of the output tokens. On AA-Briefcase, medium scored 1372 against 1578 at max while using under a quarter of the tokens.
The healthcare numbers make the tradeoff concrete. On PhysicianBench, the pass rate climbed from 17.8% at low effort to 43.0% at max, and the model's time per task went from 48 seconds to about 11 minutes.
At max, it stops being a quick model. It becomes a slow, cheap, careful one.
So pick the effort per job, not per app. A sensible starting point:
- Classification, routing and tagging: start at low and only move up if accuracy on your test set falls short.
- Summaries and extraction: start at medium, the default.
- Analysis you'd otherwise send to Sonnet: try max, and compare cost and quality against Sonnet on twenty real examples.
- Anything a person waits on in real time: stay at low or medium, because max can take minutes.
Step three is the interesting one. Max-effort Haiku might replace Sonnet on some of your analytical work at a lower price. It might not. Twenty real examples will tell you more than any benchmark table, including the ones above.
If the difference between thinking modes is fuzzy, our explainer on when a reasoning model beats a standard one lays out the tradeoff in plain terms.
The prep cook job: subagents under a bigger model
Back to the kitchen.
A good restaurant doesn't have the head chef dice onions. The chef decides the menu and plates the hard dishes. The prep cook does the hundred small, fast, repetitive jobs that make the chef's work possible.
That's the job Anthropic is pitching Haiku 5.5 for, almost word for word. The launch page says it pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work.
The customer quotes describe the same setup. Rogo, a finance AI company, says a bigger model builds the deck while a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs. Cognition says its Devin Fusion setup holds a FrontierCode score of 66.2 with Opus 5.5 as the lead and Haiku 5.5 as the sidekick, while cutting cost and latency.
Look at that number for a second. Haiku 5.5 alone scores 46.4% on FrontierCode Main. Paired under Opus, the system scores 66.2. The small model didn't get smarter. It got put in the right job.
Here's how that pattern works in a real build:
- A large model plans the task and breaks it into pieces.
- Haiku 5.5 runs the narrow pieces in parallel: read this file, pull this figure, summarize this thread, check this page.
- The large model reads the results and makes the decisions.
- A test or a human checks the output that matters most.
The failure mode is putting Haiku in the planner seat because it's cheap. The planner is the one place a mistake compounds through every step after it.
The same idea works outside code. Draft a hard strategy memo with your best model, then have the small fast one pull quotes, check dates and reformat tables. If you already switch between models mid-task, our guide on running several models in one workflow shows the mechanics.
Who should switch to Claude Haiku 5.5 this week, and who should wait
Most readers fall into one of four situations. Find yours.
| If you are | Do this | Because |
|---|---|---|
| Running Haiku 4.5 in production | Test Haiku 5.5 on a sample now and plan the switch | Better on nearly every benchmark at a fraction of the price |
| Using Sonnet for high-volume, simple jobs | Move those jobs to Haiku 5.5 at low or medium effort | Same work, a much smaller bill |
| Building a coding agent | Keep Sonnet or Opus as the lead, add Haiku 5.5 as the subagent | Anthropic says the bigger models stay better at complex coding |
| Using AI in a chat window, not the API | Keep using the model that writes best for you | Haiku's gains show up at volume, not in one conversation |
That last row is the objection most people will have, so let me say it for you.
"I don't run a million requests a month. Why do I care about a cheap model?"
Fair. If you write one email a day with AI, the price of Haiku means nothing to you, and the smartest model you can get is the right call.
But the speed still matters, and so does the computer use score. The fast model is the one that will sit inside the tools you already use, answering support chats, filling forms and clicking through web pages for you. You'll meet Haiku 5.5 even if you never pick it.

For teams with real volume, the switch is the easiest decision in this post. Asana reported over 30% lower latency on task completions and up to 2.5x faster inference per agent turn in its tests. Run your own sample first, since your prompts are not Asana's prompts.
And if you're deciding between Anthropic and OpenAI more broadly, small models are only one piece. Our Claude vs ChatGPT comparison covers how the two families fail in opposite directions.
Safety notes that matter if you build on the API
Anthropic's headline safety claim is that Haiku 5.5 shows "far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse" than Haiku 4.5. The system card backs a lot of that up.
The prompt injection numbers are the strongest. On Gray Swan's indirect prompt injection benchmark, the attack success rate at 15 attempts fell from 83.2% on Haiku 4.5 to 7.1% on Haiku 5.5. In coding environments with Anthropic's prompt injection probes on, no attack succeeded in any of 40 scenarios.
For anyone giving a small model access to an inbox or a browser, that's the number to care about. A hidden instruction in an email is how an agent gets turned against you.
The cyber safeguards are narrower than on Anthropic's bigger models. They allow a wider range of defensive security work than Sonnet 5.5's, but still block penetration testing. Security teams who hit a wall can apply to Anthropic's Cyber Verification Program.

Now the part developers need to read twice.
Anthropic's consumer app, claude.ai, wraps every model in a system prompt with extra safety language. The API doesn't. The system card says several mental health safeguards rely on that prompt, and on the raw API Haiku 5.5 helped draft suicide notes more often than Haiku 4.5, with the clearest problem when thinking was turned off.
Anthropic's instruction is direct: developers deploying Haiku 5.5, particularly with thinking disabled, should add their own safeguards. The same goes for child safety and disordered eating, where the system card asks API developers to add the protections claude.ai already has.
If you're putting Haiku 5.5 in front of the public, do three things before launch:
- Write a system prompt with explicit rules for self-harm, eating and minors, and test it with hard examples.
- Keep thinking on for any conversation that can turn personal, even at low effort.
- Route anything that reads like a crisis to a human, every time.
That's the vendor telling you where the gaps are. Listen.
How to try Claude Haiku 5.5 today
There are three practical routes, depending on how you work.
- On the Claude Platform, call the model as
claude-haiku-5-5and follow Anthropic's Haiku 5.5 migration guide if you're coming from Haiku 4.5. - On a cloud you already pay for, look for it in Amazon Bedrock, Google Cloud or Microsoft Azure, where Anthropic says it is live now.
- In a claude.ai subscription, check the model picker, and if you're on Max or Team, put the new monthly API credit toward a test run.
Before you switch anything in production, run the same test every time a model ships:
- Pull fifty real requests from last week, the boring ones included.
- Run them through your current model and Haiku 5.5 at low and medium effort.
- Count the tokens, since the new tokenizer uses slightly more per task.
- Score the answers yourself, blind if you can.
- Switch only the jobs where Haiku matched or beat what you have.
That's an afternoon of work, and it beats every table in this post, mine included.
One note on Magai, because people will ask. Magai already includes Claude models alongside GPT, Gemini, Grok and dozens more, and you can switch between them mid-conversation without losing the thread. Haiku 5.5 isn't on our list as I write this. Check the live models list, which updates from the app every day, before you assume either way.
For more on how the major models stack up, the artificial intelligence section of the blog keeps up with every release, including last week's Mistral Large 4.
So, where are you still paying a head chef to dice onions?
Hire the prep cook. Keep the chef.
More in Artificial Intelligence

Artificial Intelligence
Perplexity Pro: What $20 a Month Actually Buys
Perplexity Pro costs $20 a month or $200 a year. Here's the exact feature list, the pricing Perplexity stopped spelling out, and whether it's worth it.

Artificial Intelligence
Mistral Large 4: What It Is and Who Should Use It
Mistral Large 4 is a 1 trillion parameter open-weight model with real strength in security work. Specs, benchmarks, price, and who should try it now.

Artificial Intelligence
Grok vs ChatGPT: The Switching Cost Nobody Prices
Benchmarks say Grok and ChatGPT are close. The real difference shows up the moment you switch between them mid-task, and it costs you every time.