Anthropic has introduced Claude Haiku 5.5, its cheapest and fastest small model. It targets high-volume tasks: summarization, compaction, classification, and subagent work. The model keeps a 1M-token context window and up to 128K output tokens. Pricing starts at $0.10 per million input and $0.50 per million output tokens for prompts up to 100K. That is 90% below Haiku 4.5 on short prompts. Ready to see where it wins and how to wire it into your stack?
What exactly did Anthropic ship with Haiku 5.5?
Anthropic shipped a small Claude Haiku 5.5 tuned for speed and price at scale. It keeps a 1M context, up to 128K output tokens, and adds thinking control: adaptive thinking on by default and an effort parameter defaulting to medium.
The model accepts text and images, outputs text, and has a June 2026 knowledge cutoff. Batch jobs in beta support up to 300K output tokens. That enables long batch scenarios without redesigning your pipelines.
Two configuration caveats matter. Non-default temperature, top_p, or top_k values return a 400 error. The new tokenizer also counts roughly 30% more tokens than Haiku 4.5 for the same text. The migration guide covers both.
Is it deployable right away? Yes. Haiku 5.5 is available as a hosted API on Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and the Claude Platform on AWS. Need baseline speed, steady context, and minimal cost? This fits the brief.
How does pricing work and what will you actually pay?
Pricing follows two tiers. Up to 100K prompt tokens, input costs $0.10 and output $0.50 per 1M tokens. You also pay $0.01 per 1M for cache reads and $0.125 for 5-minute cache writes. Above 100K, rates rise to $0.50 input and $2.50 output.
For reference, Haiku 4.5 charged $1 input and $5 output. Anthropic says about 90% of Haiku 4.5 requests stayed under 100K tokens. After adjusting for the tokenizer, it estimates Haiku 5.5 to be about 75% cheaper on average. Batch processing takes another 50% off.
On short prompts, Haiku 5.5 matches GPT-6 Luna’s list price. But watch the thresholds: Haiku’s higher tier kicks in above 100K, Luna’s above 272K. If your prompts often cross the line, your cost structure will shift.
Two practical notes. The 400 error on temperature/top_p/top_k can break older integrations—check your clients. And the ~30% higher token counts mean “old” prompts may cross the 100K limit earlier than expected. Planning to batch to capture the extra 50% discount?
How does Haiku 5.5 perform on benchmarks?
By the published figures, Haiku 5.5 advances sharply over Haiku 4.5 and posts competitive results against GPT-6 Luna and others. Gains on OSWorld 2.1 (offline subset) and Terminal-Bench 4.0 stand out. Yet Sonnet 5.5 remains the family leader.
OSWorld 2.1 (offline subset) shows 72.4% for Haiku 5.5 versus 48.9% for GPT-6 Luna and 15.7% for Haiku 4.5. Terminal-Bench 4.0 reports 39.2% for Haiku 5.5 versus 16.4% for Luna and 0.0% for Haiku 4.5. FrontierCode 1.1 (Main) reads 46.4% for Haiku 5.5 versus 42.4% for Luna.
On Humanity’s Last Exam, Haiku 5.5 scores 45.9% without tools and 57.4% with tools. On GDPval-AA v2.1 it records 1620, versus 1437 for Luna and 735 for Haiku 4.5. Meanwhile, Sonnet 5.5 leads every row, including 70.6% on Terminal-Bench 4.0. For complex agentic coding, Anthropic recommends Sonnet 5.5 and Opus 5.5.
All figures below are Anthropic-reported.
What does that mean for selection? Haiku 5.5 is not a “lead coder,” but a strong subagent and high-throughput workhorse. When you need higher scores and broader agency, look to Sonnet 5.5 or Opus 5.5. Need speed, steady quality, and price? Haiku 5.5 fits.
Where does Haiku 5.5 work best right now?
Three workloads are the best fit. First, subagent roles under Opus 5.5 or Sonnet 5.5. For example, at Rogo a Haiku 5.5 subagent pulls a 10-K revenue line while the bigger model builds the deck.
Second, high-volume document Q&A and summarization. AlphaSense tested it on a feature handling about 8 million calls per week. That economics favors scaled enterprise scenarios.
Third, speed-sensitive tasks like live customer support and browser use. Here even small token savings with steady latency deliver tangible gains in cost and experience.
So how do you choose between a “subagent” and a “lead” model? Match by job: mass reductions and extractions—Haiku 5.5; long, complex planning and coding—Sonnet 5.5 or Opus 5.5. That simplifies agentic system design.
How does Haiku 5.5 stack up against its closest competitors?
On short context, list pricing is the same for Haiku 5.5 and GPT-6 Luna: $0.10 per 1M input and $0.50 per 1M output. Gemini 3.5 Flash-Lite charges $0.30 and $2.50 respectively. Long prompts change the picture: for Haiku above 100K it’s $0.50/$2.50; for Luna above 272K it’s $0.20/$0.75; Gemini uses a flat rate.
Cache reads cost $0.01 for Haiku and Luna, while Gemini charges $0.03 plus storage. Context windows are close: 1M in Haiku, 1,050,000 in Luna, and 1,048,576 in Gemini. Max output is 128K in Haiku and Luna, and 65,536 in Gemini.
By inputs, Haiku and Luna accept text and images; Gemini supports text, image, video, audio, and PDF. Reasoning control: Haiku has adaptive thinking plus effort (default medium); Luna offers reasoning.effort none to max (default medium); Gemini supports thinking. Computer use: SDK support in beta for Haiku, Responses API for Luna, and Preview for Gemini. Batch discount is 50% across all three.
Knowledge cutoffs: June 2026 for Haiku, May 18, 2026 for Luna, and not listed on the Gemini model page. Where to run? Haiku on Claude API, Bedrock, Google Cloud, Microsoft Foundry, and the Claude Platform on AWS. Luna on the OpenAI API. Gemini on the Gemini API. Pulling from the Key Takeaways: same short-context list price as Luna, while Luna is cheaper above 100K; Haiku 5.5 is positioned as a subagent under Opus 5.5 and Sonnet 5.5, not a coding lead. Ready to set your thresholds and assign the right role?
Based on the provided source.