Claude Haiku 5.5 scores 1,587 in Code Arena, about 260 points above Haiku 4.5
Arena reports Claude Haiku 5.5 at 1,587 points and rank 30 in WebDev, roughly 260 points above Haiku 4.5. Listed API prices of $0.10/$0.50 per million input/output tokens are 90% lower than its predecessor's.

TL;DR
- Haiku 5.5 (High) debuted at 1,587 points and rank 30 in Code Arena: WebDev, +257 points over Haiku 4.5, according to Arena's comparison and its debut report.
- Short-prompt API prices fell 90% to $0.10/$0.50 per million input/output tokens in Anthropic's pricing table. The company estimates 75% lower average running costs in its launch announcement.
- Cheaper tokens can still produce more expensive tasks: Vals AI found higher costs than Haiku 4.5 on every benchmark both models completed.
- The release brings a 1M-token context window, noted in the launch-day roundup, and Haiku's first adjustable effort setting, announced by Anthropic.
The WebDev leaderboard gives Haiku a ±20-point interval, wider than its six-point lead over GPT-6 Luna. Simon Willison's drawing tests ranged from a seven-second low-effort pelican to a five-minute max-effort one.
Code Arena: WebDev
Arena evaluated Haiku 5.5 at high effort and placed it just outside the cost-performance Pareto frontier.
Arena reported these gaps in points and listed token prices:
- Haiku 4.5: Haiku 5.5 scores 257 points higher at 90% lower prices.
- Sonnet 5.5 (High): 128 points lower at 95% lower prices.
- Opus 5 (High): 70 points lower at 98% lower prices.
- Fable 5.1 (Max): 157 points lower at 99% lower prices.
100k-token pricing
All input, output, and cache rates increase fivefold when the prompt exceeds 100,000 tokens under Anthropic's pricing rules. Around 90% of requests to Haiku 4.5 fell within the cheaper tier's length limit, Anthropic said.
Anthropic calculates its 75% average savings using the previous traffic mix and the updated tokenizer in the launch post's footnote. The model documentation quantifies the tokenizer change: the same text counts as approximately 30% more tokens than on Haiku 4.5.
Effort and token use
Haiku 5.5 reached 43 on the Artificial Analysis Intelligence Index, up from 17 for its predecessor, +26 points, in Artificial Analysis's testing. Its weighted output-token use, including reasoning, varied sharply by effort:
| Effort | Intelligence Index | Output tokens per task |
| --- | ---: | ---: |
| Low | 30 | ~17,000 |
| Medium | 35 | ~33,000 |
| High | 38 | ~55,000 |
| Xhigh | 41 | ~89,000 |
| Max | 43 | ~162,000 |
GPT-6 Luna at max also scored 38, using approximately 50,000 output tokens per task. Moving Haiku from xhigh to max bought 2 index points for about 1.8x the output tokens.
At launch, Artificial Analysis warned that its provisional cost estimates did not yet include the higher prices above 100k tokens.
Vals AI used maximum reasoning effort in its benchmark runs. Despite the lower token rates, every benchmark shared with Haiku 4.5 cost more on the new model.
Connections and Vibe Code Bench
- Extended Connections (
high): 14.3 → 65.7, +51.4 points over Haiku 4.5, per LechMazur. Haiku finished 3 points below GPT-6 Luna athigh. - Vibe Code Bench: 11.39% → 90.44%, +79.05 points, per Vals AI. Haiku passed every test on 21 of 50 apps, compared with 24 for Sonnet 5.5.
The Connections methodology uses 940 puzzles with up to four extra decoy words. Its score averages (exact_groups / 4)² × 100, rather than counting fully solved puzzles.
Browser Use Bench
Haiku delivered roughly GPT-6 Luna-level performance at 2.8x the estimated task cost in Browser Use Bench v2.1. Cached tokens count toward the 100k threshold, which these long-running tasks frequently crossed.
The comparison covered 180 shared tasks, with output caps of 32k for Haiku and 128k for Luna. Increasing Haiku's effort from xhigh to max reduced its score from 69.5 to 65.3, -4.2 points.
In a smaller browser QA run, Haiku found four of five planted defects plus an unplanted invoice bug, costing $0.30 against Opus's $1.43, according to kevinkern's comparison.
Haiku subagents
Multiple Haiku subagents streamed drawing commands to a shared canvas in a map-building demo.
Claude Code can pair Haiku with an Opus advisor through {"model":"haiku","advisorModel":"opus"}, as daniel_mac8 demonstrated. The poster said he hadn't tested the comparison against vanilla Opus.
Day-one access and API defaults
- Cloud platforms: AWS, Google Cloud, and Microsoft Azure are available through the official rollout.
- Genspark: AI Chat, Code Agent, and Claw support Haiku 5.5, according to Genspark's announcement.
- OpenCode: The model is available in OpenCode and Go, per OpenCode.
- GitHub Copilot: Generally available in VS Code, according to the VS Code announcement.
The Claude API model ID is claude-haiku-5-5, with up to 128k output tokens per synchronous response and default effort medium, according to the model specifications. Non-default values for temperature, top_p, or top_k return a 400 error.