🚀 Released October 7, 2026 — the fastest and cheapest model of the Claude 5.5 family, far below Haiku 4.5 ($1/$5)

Claude Haiku 5.5

The fastest and cheapest model in Anthropic's Claude 5.5 family. The official docs describe it as built for “high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks.” Pricing is $0.10 input / $0.50 output for prompts up to 100,000 tokens — and $0.50 / $2.50 once a prompt goes above 100,000 tokens.

$0.10
Input / 1M (≤100k)
$0.50
Output / 1M (≤100k)
1M
Context
Fastest
Latency (official table)

The short version

  • One tenth of Haiku 4.5 on input and output. $0.10 / $0.50 up to 100,000 tokens, against $1.00 / $5.00 on Haiku 4.5.
  • Above 100,000 tokens the rates jump 5x to $0.50 input / $2.50 output, so the official table reads “From $0.10 / $0.50”. The 1M window is real, but the price of filling it is not the headline rate.
  • Cache reads are 10% of input: $0.01 (and $0.05 above 100k). Cache writes are $0.125 for the 5-minute TTL and $0.20 for one hour.
  • Thinking is adaptive and on by default, with a default effort of medium — more conservative than Sonnet 5.5's high.
  • The tokenizer changed. Verbatim: it “uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5.” Compare cost with token counts in hand.
  • temperature, top_p and top_k can no longer be set to non-default values — they return a 400 error.

Pricing (USD per 1M tokens, from the official docs)

Values as published on the Haiku 5.5 overview page and the official pricing documentation; the Haiku 4.5 column matches our pricing comparison.

ItemHaiku 5.5 (≤100k tokens)Haiku 5.5 (over 100k)Haiku 4.5 (previous gen)
Input$0.10$0.50$1.00
Output$0.50$2.50$5.00
Cache write (5-minute TTL)$0.125$0.625$1.25
Cache write (1-hour TTL)$0.20$1.00$2.00
Cache read (cache hit)$0.01$0.05$0.10
Batch API50% off input and output50% off

※ Haiku 5.5 figures are as published. The Haiku 4.5 cache figures are calculated by us from the official multipliers (1.25x for 5-minute writes, 2x for one hour, 0.1x for reads), verified 2026-10-09.

Source: Claude Haiku 5.5 — overview (Pricing / Availability), checked 2026-10-09.

Where the two-tier pricing actually bites

The rate switches the moment a single prompt crosses 100,000 tokens. Worked examples:

CaseHaiku 5.5Claude Sonnet 5.5
200,000 input + 10,000 output tokens$0.125
($0.50/M × 0.2 + $2.50/M × 0.01)
$0.50
($2/M × 0.2 + $10/M × 0.01)
50,000 input + 5,000 output tokens$0.0075
($0.10/M × 0.05 + $0.50/M × 0.005)
$0.15
($2/M × 0.05 + $10/M × 0.005)
  • Up to 100k the rate is 1/20 of Sonnet 5.5. For high-volume classification, routing and extraction, that is an order-of-magnitude difference.
  • Above 100k it is still about a quarter of Sonnet 5.5. The honest framing is not “expensive above 100k” but “unusually cheap below it”.
  • Reused long prefixes belong in the cache: $0.01/M read (against $0.10/M on Sonnet 5.5) — the pattern that fits RAG and agent system prompts.

Four things to watch when migrating from Haiku 4.5

1. About 30% more tokens
The same text counts as roughly 30% more tokens on the new tokenizer. A 10x lower unit price works out to about 7.7x cheaper in practice (1 ÷ (0.1 × 1.3)) — not 10x.
2. Thinking is on by default
Adaptive thinking, default effort medium. Budget for extra output tokens, and steer depth with effort.
3. No temperature parameters
Verbatim: “Omit temperature, top_p, and top_k, since a non-default value for any of them returns a 400 error.”
4. Thinking blocks are account-bound
Verbatim: “Its thinking blocks work only in the account that produced them, or in an account linked to it.” Multi-account setups cannot carry them across.

Specifications and model IDs

ItemValue (official page)
Context window1M tokens
Max output128K tokens (300K on the Batch API with the output-300k-2026-03-24 beta header)
ThinkingAdaptive (on by default); default effort medium
LatencyFastest (official comparison table)
Modalitiestext, images → text
Knowledge cutoffJun 2026 (training data cutoff Jun 2026)
StatusActive (latest). Retirement not sooner than October 7, 2027.
PlatformModel ID
Claude APIclaude-haiku-5-5
Amazon Bedrockanthropic.claude-haiku-5-5
Google Cloud (Vertex AI)claude-haiku-5-5
Microsoft Foundryclaude-haiku-5-5
Claude Platform on AWSclaude-haiku-5-5

What the docs say, verbatim

“Claude Haiku 5.5 is built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. It supports adaptive thinking with the effort parameter, a 1M token context window, and up to 128k output tokens. It uses the same newer tokenizer as Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Claude Haiku 4.5.” — Claude Platform Docs, “Claude Haiku 5.5 — overview” (checked 2026-10-09)

Pricing, specifications and model IDs above are as published. We do not fill gaps with estimates: where the vendor publishes no benchmark table, we say so.

Good fit, poor fit

Good: bulk classification, extraction, routing
Many short calls. Up to 100k tokens the input rate is $0.10/M against $2/M on Sonnet 5.5.
Good: subagents and preprocessing
The docs name subagent tasks explicitly — a natural front stage for a stronger model.
Good: latency-sensitive chat
Fastest latency in the official comparison table.
Poor: audit or contract decisions
Retention, operator access and training terms are not on the model page. Start from ZOA, ZDR and the four requirements.
Poor: pipelines that set temperature
Non-default temperature, top_p or top_k returns a 400 error.
Reconsider: 100k+ single-shot prompts
Rates become $0.50 / $2.50. For whole-document work, compare the total against Sonnet 5.5.

Related pages

⚠️ Disclaimer

  • Prices and specifications were checked on 2026-10-09 against Claude Platform Docs (Haiku 5.5 overview and the official pricing documentation). Vendors change them without notice.
  • Haiku 4.5 cache figures are our own calculation from the official multipliers.
  • This site is informational and does not endorse or represent any provider.