Anthropic released Claude Haiku 5.5 on October 7 at $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens. Haiku 4.5 cost $1 and $5. Anthropic calls that 90% lower for the short-prompt case and 75% lower on average, and the headlines have repeated the 90. The number is real and the unit is wrong. Nobody ships tokens. They ship finished tasks, and a task’s cost is the token price times how many tokens the model chooses to spend times how often it has to try again. My read is that Haiku 5.5 moves the first factor a long way and hands you control of the second, and the builders who profit will be the ones who price the task.
The 90% has three asterisks
The first is the cliff. Haiku 5.5 is the one current model whose price depends on prompt length: at 100,000 tokens and below it is $0.10 in and $0.50 out, and above that it is $0.50 in and $2.50 out, according to the pricing page. Every other model at 4.6 or later gets the full 1M window at one rate. Anthropic’s own launch page concedes the consequence, listing the long-prompt cut as 50%, not 90%. A 99,000-token prompt costs about one cent of input on Haiku 5.5 and about twenty cents on Sonnet 5.5 at its flat $2. A 150,000-token prompt costs about 7.5 cents against thirty. Haiku is still cheaper in both cases, but the ratio falls from twenty to one to four to one at the line, so an agent whose context grows past 100,000 tokens mid-run changes its economics without anyone touching the config.

The second is the tokenizer. The same pricing page says Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text, with the exact figure depending on content. Haiku 4.5 predates that, and Haiku 5.5 does not. The page does not say the comparison explicitly, so this is my inference: $0.10 per token of the new kind buys about what $0.13 buys of the old kind, which puts the like-for-like cut near 87% rather than 90%. That is a rounding error against a tenfold drop, and it is also the sort of thing a spreadsheet built from sticker prices will miss.
The third is that the benchmark table does not say what it cost to produce. Which brings up thinking.
Thinking is a line item you now control
Haiku 5.5 is the first Haiku with effort levels, and it thinks by default. Anthropic’s prompting guide says thinking counts toward max_tokens, that a limit sized for Haiku 4.5 requests that ran without thinking can cut the reply off, and that telling the model in the prompt to answer directly did not stop it from thinking in Anthropic’s testing. The lever is the effort setting, and the effort documentation puts the default at medium. I have not found a billing line that exempts thinking from output pricing, so I am treating those tokens as output tokens at $0.50 per million. That is an inference about billing, not a quoted rule.
The same guide is candid about what the cheap setting costs you. In a long agent prompt, low effort makes the model more likely to skip a search, stop early, or skip a check. Anthropic measured that moving from low to medium roughly halved early stopping and more than doubled output tokens per attempt, and it notes that at low and medium effort the model sometimes reports a code change as finished without running anything that exercises it. Anthropic ships the prompt text that mostly fixes this, at the price of more tokens. So the honest cost of a Haiku 5.5 task is not a constant. It is a curve you choose a point on, and the curve is steeper than the sticker suggests.
Here is the strongest fact against my own thesis, and I would rather say it than have you find it. Even with that curve, the arithmetic still favors the small model by a wide margin. Output is $0.50 per million on Haiku 5.5 and $10 on Sonnet 5.5, a factor of twenty. A Haiku run that burns five times the tokens of a Sonnet run, retries included, still costs a quarter as much. If your task is one where Haiku succeeds at all, effort-driven token growth will not erase the saving. The danger is not that Haiku gets expensive. It is that it quietly fails the check you forgot to write, at a price low enough that nobody audits it.
One operational detail belongs here because it eats the other saving. The effort documentation says changing the top-level effort value between requests invalidates the prompt cache for that conversation, and the workaround, per-message effort, is a beta that needs a header and adaptive thinking. A router that dials effort up and down per step on one long conversation can pay for the flexibility in cache misses. Anthropic did halve Sonnet 5.5’s cache-read price, from $0.20 to $0.10 per million the same day, which makes the cache worth protecting.
Where the gap shows
Anthropic’s launch table compares Haiku 5.5 with Haiku 4.5 and Sonnet 5.5, and I would read the Sonnet column first. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% against 70.6% for Sonnet 5.5. On OSWorld 2.1, the offline subset, it scores 72.4% against 83.9%. On GDPval-AA v2.1 it is 1,620 Elo against 1,840, on AA-Briefcase v1.1 it is 1,578 against 1,824, on Humanity’s Last Exam it is 45.9% against 56.9% without tools and 57.4% against 64.5% with them, and on Chartography without tools it is 46.4% against 61.6%. These are vendor numbers. The page does not define what the offline subset is, and it does not state the effort level behind each row. VentureBeat reports that the roughly 39% Terminal-Bench figure comes at maximum effort and that medium lands near 20%. I could not confirm that split in Anthropic’s own text, so treat it as a lead to check, not a fact. If it holds, the headline score and the default setting are different products.
What the table does establish is the size of the jump. Haiku 4.5 scored 0.0% on Terminal-Bench and 15.7% on OSWorld. A model that could not complete one terminal task in the suite is now at 39%, and one that managed 15.7% on computer use is at 72.4%. That is a change in what the small tier is allowed to attempt, not an increment. The customer quotes are weaker evidence. HubSpot reports 92.8% on its own CRM suite, the best it has seen, and Box reports eleven points over Haiku 4.5 at about half the latency, with no scale given for the points. Asana reports a drop of more than 30% in latency and 2.5x faster inference per agent turn. They are all real data points about someone else’s workload.
VentureBeat also notes that OpenAI’s GPT-6 Luna matches the $0.10 and $0.50 entry price, with its surcharge starting at 272,000 input tokens instead of 100,000. If the long-context cliff matters to you, that is a pricing difference worth the same scrutiny as the benchmark. VentureBeat’s Luna comparison also comes from vendor-reported results, and it favors Haiku: 1,620 against 1,437 on GDPval-AA and 39.2% against 16.4% on Terminal-Bench.

Route by whether failure is visible
Anthropic’s models overview describes Haiku 5.5 as the model for high-volume, latency-sensitive tasks such as classification, extraction and routing. Take that literally and add one rule. Give the small model the steps where a wrong answer is cheap to detect: output that has to validate against a schema, a label checked against a known set, a duplicate check, a triage that a downstream step can overturn. Keep the larger model on the steps where a wrong answer looks identical to a right one until a reader finds it. The writing step of an automated publication belongs to the second kind. A schema validator catches a missing field. Nothing catches a confident sentence that is wrong.
Two practical things follow. First, build the refusal branch before you route anything. Haiku 5.5 runs safeguard classifiers that can return stop_reason: "refusal" in four categories, and the guide says that for people moving from Haiku 4.5 these refusals are new, that there is no server-side fallback, and that resending the same request usually returns another refusal. A router that treats the cheap tier as a drop-in will fail on whatever tripped the classifier. Second, measure cost per accepted output instead of cost per call. Count the retries, the thinking tokens, the cache misses, and the humans who have to look at the rejects, then compare that number across tiers on your own data. Anthropic says as much, advising that you compare xhigh and max against Sonnet 5.5 on performance, cost and speed before committing.
I have not run Haiku 5.5 through my own pipeline yet, so this is analysis of the documents, not a result. The first test I would run is the dull one: take the cheapest step you currently give a frontier model, run a few hundred real inputs through Haiku at low, medium and high effort, and count the failures your checker catches and the ones it does not. The price drop is genuine and the small tier can do real work now. Pricing it per token is how you find out about the rest from your invoice.

AI-generated editorial illustration · TemperatureZero · October 8, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive