Anthropic Launches Claude Haiku 5.5 to Slash API Costs
Anthropic has launched Claude Haiku 5.5, a highly efficient small model designed to dramatically lower the cost of high-volume tasks like summarization and subagent orchestration.

Anthropic has launched Claude Haiku 5.5, its most economical and rapid small model yet, featuring a 1-million-token context window, up to 128,000 output tokens, and a June 2026 knowledge cutoff. Generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, it supports text and image inputs. For prompts up to 100,000 tokens, pricing is $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens, rates rise to $0.50 input and $2.50 output. Cache reads cost $0.01 and five-minute cache writes cost $0.125 per million tokens.
This pricing is 90 percent cheaper than Claude Haiku 4.5 (which charged $1 input and $5 output) for prompts under 100,000 tokens, where 90 percent of Haiku 4.5 requests fell. Although a new tokenizer counts text as roughly 30 percent more tokens, Anthropic estimates Haiku 5.5 is still 75 percent cheaper on average. Beta batch processing, supporting up to 300,000 output tokens, cuts costs by 50 percent. Developers must avoid non-default temperature, top_p, or top_k values, which return a 400 error. Meanwhile, GPT-6 Luna has identical short-context rates but is cheaper for a 150,000-token prompt because its higher tier ($0.20 input, $0.75 output) only triggers above 272,000 tokens. Gemini 3.5 Flash-Lite charges a flat $0.30 input and $2.50 output, with a $0.03 cache read plus storage, a 1,048,576-token context, and a 65,536-token output limit.
Haiku 5.5 introduces an adjustable effort setting for adaptive thinking, defaulting to medium. On OSWorld 2.1, it scored 72.4 percent, beating GPT-6 Luna's 48.9 percent and Haiku 4.5's 15.7 percent. On Terminal-Bench 4.0, it scored 39.2 percent, outperforming Luna's 16.4 percent and Haiku 4.5's 0.0 percent. It also achieved 46.4 percent on FrontierCode 1.1 (versus 42.4 percent for Luna), 1620 on GDPval-AA v2.1 (versus 1437 for Luna and 735 for Haiku 4.5), and 45.9 percent on Humanity’s Last Exam without tools (57.4 percent with tools).
For practitioners, Sonnet 5.5 still leads benchmarks, including 70.6 percent on Terminal-Bench 4.0. Anthropic recommends Haiku 5.5 as a subagent under Sonnet 5.5 or Opus 5.5. At Rogo, a Haiku 5.5 subagent pulls 10-K revenue lines while a larger model builds decks. AlphaSense tested it on a feature handling 8 million calls weekly. It is also optimized for speed-sensitive tasks like live customer support and browser use.
This is our own summary of reporting by MarkTechPost



