Claude Haiku 5.5 Review: Ten Cents a Million Tokens, and Actually Good
Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output, 90% less than Haiku 4.5, and it's genuinely good. For API work on a tight budget, it may be the best model at its price right now. Just watch the 100K-token price cliff.
- AI
- Claude
- LLMs

I wasn't a fan of Sonnet 5.5. In my review I called it pretty good, said it wasn't my pick for hard coding, and spent a whole section complaining about its cache pricing.
Nine days after Sonnet, Anthropic released Claude Haiku 5.5. I wasn't expecting much from it. Small models are usually the ones you put up with, not the ones you get excited about.
This one is different. Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output, 90% less than Haiku 4.5, and it's genuinely good. On Artificial Analysis's independent testing, at xhigh effort it matches Sonnet 5.5 at medium for about a quarter of the cost per task.
So here's my take, up front: for API work on a tight budget, Haiku 5.5 may be the best model you can buy at its price right now. It has two asterisks, and I'll get to both. It also completes Anthropic's 5.5 lineup, and the same launch fixed the Sonnet complaint I made at the end of September.
Haiku 5.5 at a glance
- API price, input / output
- $0.10 / $0.50
- Per million tokens, for prompts up to 100K. 90% below Haiku 4.5
- Intelligence Index, max effort
- 43
- 26 points above Haiku 4.5, and ahead of GPT-6 Luna's 38
- Cost per task, xhigh effort
- $0.12
- The same score as Sonnet 5.5 at medium (41), for a quarter of its $0.48
- Price over 100K prompt tokens
- 5x
- $0.50 / $2.50, on the whole request, cache reads included
Released October 7, 2026 as claude-haiku-5-5, with a 1M-token context window and up to 128K tokens of output. Index score and cost per task from Artificial Analysis, which doesn't yet apply the 100K price step; prices from Anthropic.
What Anthropic shipped
Haiku 5.5 came out on October 7 as claude-haiku-5-5. It's the third and last model in the Claude 5.5 family, after Opus 5.5 on September 22 and Sonnet 5.5 on September 28.
The upgrades over Haiku 4.5 are bigger than usual for a small model:
- A 1M-token context window, up from 200K, and up to 128K tokens of output, up from 64K.
- Effort settings, from low to max. It's the first Haiku to have them. Thinking is on by default, and the default effort is medium.
- Price: $0.10 / $0.50 per million tokens for prompts up to 100K, the same as OpenAI's GPT-6 Luna, the smallest model in the GPT-6 family I covered in my GPT-6.1 Sol post.
Anthropic pitches it as "the cheapest, fastest, and most capable small model we've ever released." It's aimed at high-volume work: summaries, classification, database queries, compacting long conversations, and acting as a sub-agent that Opus or Sonnet hands smaller jobs to. Anthropic also suggests it for live customer support and browser automation, because it's the fastest model it has at standard speed.
Anthropic is also upfront about the limit: for complex agentic coding, it says, "Sonnet 5.5 and Opus 5.5 remain better choices." I appreciate that. It matches what the numbers show.
It's available on Anthropic's API, AWS, Google Cloud and Microsoft's cloud.
The price: 90% off, with two asterisks
Here's the full price sheet next to Haiku 4.5 and Luna, from Anthropic's pricing page and OpenAI's Luna page:
| Per million tokens | Haiku 4.5 | Haiku 5.5, prompt up to 100K | Haiku 5.5, prompt over 100K | GPT-6 Luna, up to 272K |
|---|---|---|---|---|
| Input | $1.00 | $0.10 | $0.50 | $0.10 |
| Output | $5.00 | $0.50 | $2.50 | $0.50 |
| Cache read | $0.10 | $0.01 | $0.05 | $0.01 |
| Cache write (5 min) | $1.25 | $0.125 | $0.625 | $0.125 |
| Batch input / output | $0.50 / $2.50 | $0.05 / $0.25 | $0.25 / $1.25 | Half of standard |
A tenth of Haiku 4.5's price, until the prompt passes 100K tokens
API price per million output tokens. Input costs a fifth of output on every row, so the picture is the same for input.
- Sonnet 5.5
- Haiku 4.5
- Haiku 5.5
- GPT-6 Luna
Sonnet 5.5
Haiku 4.5
Haiku 5.5, prompt over 100K
Haiku 5.5, prompt up to 100K
GPT-6 Luna, up to 272K
Source: Anthropic and OpenAI API pricing, as of October 11, 2026. Standard rates; batch processing halves all of them.
Show the data
| Model | Input per 1M tokens | Output per 1M tokens |
|---|---|---|
| Sonnet 5.5 | $2.00 | $10.00 |
| Haiku 4.5 | $1.00 | $5.00 |
| Haiku 5.5, prompt over 100K | $0.50 | $2.50 |
| Haiku 5.5, prompt up to 100K | $0.10 | $0.50 |
| GPT-6 Luna, up to 272K | $0.10 | $0.50 |
On paper, that's the whole story: a tenth of the price, tied with Luna. Here are the two asterisks.
Asterisk one: the 100K cliff
When a prompt goes over 100,000 tokens, the whole request is billed at the higher rate, five times the normal price. Anthropic's docs are specific about what counts: every input token in the request, including cache reads and cache writes. Each request is priced on its own.
For a typical classification or extraction job, you'll never notice. Anthropic says about 90% of Haiku 4.5's requests were under 100K tokens.
Agent loops are where it bites. An agent re-sends its whole history on every turn, so its prompts keep growing, and once they cross 100K every later call costs five times as much. Luna's long-context surcharge starts much later, at 272K, and it's smaller: double the input price and 1.5 times the output price.
A quick example from my own arithmetic: a 150K-token prompt with a 2K-token answer costs about $0.08 on Haiku 5.5. The same request on Luna costs about $0.016, because it's still under Luna's line.
To be fair to Haiku, even over the line its input and output prices are a quarter of Sonnet 5.5's. It's a cliff, not a disaster. But you should know where it is.
Asterisk two: the tokenizer
Haiku 5.5 uses Anthropic's newer tokenizer, the one Sonnet 5.5 and Opus 5.5 use. Per Anthropic, it produces about 30% more tokens for the same text than the older one Haiku 4.5 used.
So "90% cheaper per token" is more like 87% cheaper for the same text. Anthropic's own estimate, which also accounts for the share of requests above 100K, is that the average workload costs about 75% less than on Haiku 4.5. That's still a huge cut. It's just not the number in the headline.
Why I'm impressed: it's good, not just cheap
Cheap small models aren't new. What's new is a cheap small model that's this capable.
On Artificial Analysis's Intelligence Index, Haiku 5.5 scores 43 at max effort. That's 26 points above Haiku 4.5, and ahead of the other small models Artificial Analysis compares it with: GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38). Just above it sits Kimi K3 at 44, a 2.8-trillion-parameter open-weights model that is anything but small.
Anthropic's launch results tell the same story:
A different class of model from Haiku 4.5
Anthropic's launch results. Haiku 5.5 beats GPT-6 Luna on every test both were run on, and sits below Sonnet 5.5 on all of them.
- Haiku 4.5
- Haiku 5.5
- GPT-6 Luna
- Sonnet 5.5
OSWorld 2.1 (offline)
Terminal-Bench 4.0
FrontierCode 1.1
Chartography
Humanity's Last Exam
Source: Anthropic's Haiku 5.5 launch post. Chartography and Humanity's Last Exam are without tools; Sonnet 5.5's FrontierCode score is at xhigh effort. Anthropic didn't report Haiku 4.5 on FrontierCode or Luna on Humanity's Last Exam.
Show the data
| Benchmark | Haiku 4.5 | Haiku 5.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| OSWorld 2.1 (offline) | 15.7% | 72.4% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 0% | 39.2% | 16.4% | 70.6% |
| FrontierCode 1.1 | Not reported | 46.4% | 42.4% | 52.1% |
| Chartography | 6.4% | 46.4% | 29.1% | 61.6% |
| Humanity's Last Exam | 10.2% | 45.9% | Not reported | 56.9% |
The jump from Haiku 4.5 is enormous. On Terminal-Bench 4.0, a test of agents working in a terminal, Haiku 4.5 scored 0% and Haiku 5.5 scored 39.2%. On OSWorld, a computer-use test, it went from 15.7% to 72.4%. On GDPval-AA, a knowledge-work benchmark rated like chess, Haiku 5.5 rates 1620, against 1437 for Luna and 1840 for Sonnet 5.5.
The usual caveat applies: those are Anthropic's own numbers. Two things make me trust them more than usual:
- Artificial Analysis's own Terminal-Bench run agrees on the ranking: 33% for Haiku 5.5, against 20% for Gemini 3.8 Flash and 13% for Luna. Lower than Anthropic's figure, but the same order.
- It makes things up less. On Artificial Analysis's hallucination test, Haiku 5.5 hallucinated 40% of the time when it didn't know an answer. Gemini 3.8 Flash did 55% and Luna 77%. Haiku actually knows less (36% accuracy, against 44% for Luna and 55% for Gemini 3.8 Flash), but it's more willing to say so. For a model you run unattended a million times, I'll take that trade.
One thing to know before you quote the 39.2%: according to VentureBeat's reading of Anthropic's charts, that's at max effort. At the default medium setting, Haiku 5.5 scores about 20%.
The comparison that sold me: Haiku 5.5 vs Sonnet 5.5
This is the chart that changed my mind about this model:
Haiku 5.5 at xhigh matches Sonnet 5.5 at medium, for a quarter of the cost
Intelligence Index score against the cost of one task. Each dot is an effort level, from low on the left to max on the right, so higher and further left is better value. Sonnet 5.5 keeps climbing long after Haiku 5.5 stops.
- Sonnet 5.5
- Haiku 5.5
- GPT-6 Luna
Intelligence Index
Cost per task
Source: Artificial Analysis, as of October 11, 2026. The cost axis is logarithmic. Haiku 5.5's costs don't yet include its higher price for prompts over 100K tokens, and GPT-6 Luna is only published at max effort.
Show the data
| Model | Effort | Cost per task | Index score |
|---|---|---|---|
| Sonnet 5.5 | Low | $0.35 | 36 |
| Sonnet 5.5 | Medium | $0.48 | 41 |
| Sonnet 5.5 | High | $0.88 | 47 |
| Sonnet 5.5 | Xhigh | $2.01 | 52 |
| Sonnet 5.5 | Max | $5.46 | 56 |
| Haiku 5.5 | Low | $0.02 | 29 |
| Haiku 5.5 | Medium | $0.05 | 34 |
| Haiku 5.5 | High | $0.08 | 38 |
| Haiku 5.5 | Xhigh | $0.12 | 41 |
| Haiku 5.5 | Max | $0.21 | 43 |
| GPT-6 Luna | Max | $0.07 | 38 |
- At xhigh effort, Haiku 5.5 scores 41 for $0.12 a task. Sonnet 5.5 scores the same 41 at medium, for $0.48. Same result, a quarter of the cost.
- At max, Haiku 5.5 scores 43 for $0.21, higher than Sonnet 5.5 at medium for less than half the price.
- Against Luna, it's close to a tie. Luna at max scores 38 for $0.07. Haiku 5.5 at high scores 38 for $0.08. Above that, Haiku keeps going and Luna doesn't.
Sonnet 5.5 isn't obsolete. It climbs all the way to 56, and Haiku 5.5 tops out at 43. For hard work, the bigger models are still clearly smarter. But a lot of API work isn't hard work. It's the same well-defined task, run thousands of times. For that, paying Sonnet prices for Sonnet-at-medium quality no longer makes much sense.
One important caveat: Artificial Analysis says its Haiku 5.5 costs don't yet include the higher price for prompts over 100K tokens. Short tasks are priced right. Long ones will cost more than the chart shows.
It's fast, too. Artificial Analysis measured about 240 output tokens per second at max effort, against 137 for Luna at max. Asana, one of Anthropic's launch customers, reported task latency down more than 30%.
The other good news: Anthropic fixed my Sonnet complaint
In my Sonnet 5.5 review, my biggest criticism was that Anthropic cut cache-read prices on Opus 5.5 but not on Sonnet. Sonnet is the model people run inside long agent loops, which spend most of their money re-reading cached context.
In the same Haiku announcement, Anthropic halved Sonnet 5.5's cache-read price, from $0.20 to $0.10 per million tokens. That's 5% of the input price, the same ratio as Opus 5.5, and level with the $0.10 cached input I pointed out on GPT-6.1 Sol. Anthropic says it makes most agentic work on Sonnet about 20% cheaper.
The independent numbers agree. When I reviewed Sonnet 5.5, Artificial Analysis had it at $0.59 a task at medium effort. It's now $0.48, with the same score.
Here's the illustrative agent task from that review again: a loop that reads 2M cached tokens, sends 100K fresh input tokens and writes 50K output tokens. Now with Haiku 5.5 added:
One agent task, priced on each model
An illustrative agent loop that reads 2M cached tokens, sends 100K fresh input tokens and writes 50K output tokens. Halving Sonnet's cache price makes it exactly half of Opus here. Haiku 5.5 is cheaper again by a wide margin, even above its 100K line.
Opus 5.5
Sonnet 5.5, before Oct 7
Sonnet 5.5, now
Haiku 5.5, prompts over 100K
Haiku 5.5, prompts up to 100K
An illustrative example priced at Anthropic's API rates, not a benchmark. It assumes every model needs the same tokens to finish, and leaves out cache writes. Sonnet 5.5's cache reads fell from $0.20 to $0.10 per million tokens on October 7.
Show the data
| Scenario | Cache reads | Fresh input | Output | Total |
|---|---|---|---|---|
| Opus 5.5 | $0.40 | $0.40 | $1.00 | $1.80 |
| Sonnet 5.5, before Oct 7 | $0.40 | $0.20 | $0.50 | $1.10 |
| Sonnet 5.5, now | $0.20 | $0.20 | $0.50 | $0.90 |
| Haiku 5.5, prompts over 100K | $0.10 | $0.05 | $0.125 | $0.275 |
| Haiku 5.5, prompts up to 100K | $0.02 | $0.01 | $0.025 | $0.055 |
Sonnet 5.5 is now exactly half the price of Opus 5.5 on this task, which is what the price sheet always implied. Haiku 5.5 is in a different league: $0.055 if every request stays under 100K tokens, $0.275 if they all go over.
That's also why the 100K line matters so much. A loop that reads 2M cached tokens is, for example, 20 turns that each re-read about 100K tokens. That's right on Haiku's line. Which side of it you land on decides whether the task costs about 3% of the Opus price or about 15%.
The chart assumes every model needs the same number of tokens to finish, which won't always be true. A smaller model may need more turns, or get the job wrong. That's the junior engineer problem I wrote about when Sonnet 5 launched. Still, the gap is so wide that Haiku has a lot of room for error.
Where it falls short
It's a strong release, but it isn't perfect:
- Hard coding. Anthropic says so itself, and the benchmarks agree: 39.2% on Terminal-Bench against Sonnet 5.5's 70.6%. Use it as a sub-agent for coding, not as the lead.
- Max effort is wordy. Artificial Analysis calls it "very verbose." At max it writes about 162K output tokens per task, roughly three times as many as Luna. Going from xhigh to max adds 2 points for about 75% more cost per task.
- The first token can take a while. Thinking is on by default. On Artificial Analysis's benchmark tasks, the first token took about 10 seconds at low effort and 14 at medium. That measures hard benchmark prompts, not chat, but if you're building live support on it, test on your own prompts.
- Stricter safeguards. Its cyber safeguards are stricter than Haiku 4.5's and block penetration testing under the standard settings. Artificial Analysis's AutomationBench score (35%, against 53–60% for Luna, Gemini 3.8 Flash and GLM-5.3 Flash) was hurt by over-refusals in pre-release testing. Artificial Analysis expects the score to go up when it re-runs the test.
- Factual knowledge. As above, it knows less than Luna or Gemini 3.8 Flash. It's more honest about that, but for knowledge-heavy questions I'd still check the answer.
How I'd use Haiku 5.5
- Make it the default for high-volume API work: classification, extraction, summaries, routing, tagging. Start at low or medium effort.
- Use xhigh when quality matters. That's where it matches Sonnet 5.5 at medium. Skip max.
- Use it as a sub-agent for search, reading and retrieval, with Opus 5.5 or Sonnet 5.5 in charge.
- Keep every request under 100K tokens. Compact long conversations, start sub-agents fresh, and log
usage.input_tokens(plus cache reads) so you can see when a route goes over the line. - Re-count your tokens when moving from Haiku 4.5. The new tokenizer means the same prompts will count as more tokens.
- For hard coding and design work, use Sonnet 5.5 or Opus 5.5.
- If you're on Max or Team, Anthropic is adding monthly API credits: $100 on Max 5x, $200 on Max 20x, and up to $500 for a team. At Haiku's prices, $100 buys a billion input tokens. OpenAI went the other way at DevDay and halved the usage on its $200 Pro plan.
The lineup is complete
In 15 days, Anthropic shipped Opus 5.5, Sonnet 5.5 and Haiku 5.5. Every 5.5 model is cheaper than the model it replaces, or costs the same, and Haiku got by far the biggest cut. That's great for anyone paying the bills. It's also a bold move from a company that, as I wrote in my post on the state of the AI industry, lost more than it made last year. I hope these prices hold.
| Model | Status | Input / output per 1M tokens | Cache read |
|---|---|---|---|
| Fable 5.1 | Still on 5.1 | $10 / $50 | $0.25 |
| Opus 5.5 | Released September 22 | $4 / $20 | $0.20 |
| Sonnet 5.5 | Released September 28 | $2 / $10 | $0.10 |
| Haiku 5.5 | Released October 7 | $0.10 / $0.50 (up to 100K) | $0.01 |
That leaves Fable, Anthropic's top model, on 5.1. People online have been guessing about a Fable 5.5 for a while. As of today, Anthropic hasn't announced one, and its pricing page still lists Fable 5.1 as its top model. I'm excited to see what it does next. Opus, Sonnet and Haiku all got cheaper and better in this round, and it would be great to see Fable get the same treatment.
The verdict
Haiku 5.5 is the most impressive model in the 5.5 lineup for me. That isn't because it's the smartest; it's the least smart of the three. It's because it makes good-enough intelligence so cheap that you stop thinking about cost for most everyday API work.
It isn't flawless. The 100K cliff is a real trap for agent loops, the new tokenizer quietly takes back some of the discount, and you shouldn't hand it your hardest coding problems. But for classification, extraction, summaries and sub-agent work on a tight budget, I don't think anything at this price beats it right now.
Ten cents used to buy you a toy model. Now it buys you a good one.
A note on sources: prices come from Anthropic's pricing page and OpenAI's GPT-6 Luna page, as of October 11, 2026. Benchmarks labelled as Anthropic's, the customer results and the Sonnet 5.5 cache cut come from Anthropic's Haiku 5.5 launch post. Index scores, cost per task, speed, token counts, Terminal-Bench, hallucination and AutomationBench results come from Artificial Analysis and its model pages. The medium-effort Terminal-Bench figure is from VentureBeat, and release dates are confirmed by SiliconANGLE. The 150K-token example and the agent-task chart are my own arithmetic at list prices. Everything written as "I think" is my own opinion.