Haiku 5.5 is the stronger all-rounder.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Haiku 5.5 pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+47points ahead
- Long, multi-file coding tasks in real codebasesSWE-bench Pro+25.3points ahead
- Fixes real GitHub issues in many programming languagesSWE-bench Multilingual+16.3points ahead
Where Claude Haiku 4.5 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingHaiku 5.5
- SWE-bench ProHaiku 5.5 by 25.339.564.8+25.3
- SWE-bench MultilingualHaiku 5.5 by 16.367.483.7+16.3
Other results21 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDPval-AA 2.1Haiku 5.5 by 885 rating points7351620+885 rating
- OSWorld 2.1Haiku 5.5 by 56.715.772.4+56.7
- OSWorld 2.1 (offline subset)Haiku 5.5 by 56.315.772+56.3
- Terminal-Bench 4.0Haiku 5.5 by 39.2039.2+39.2
- HealthBench Professional length-adjustedHaiku 5.5 by 32.632.264.8+32.6
- Suicide and self-harm multi-turn appropriate response rate (API, without a system prompt)Haiku 5.5 by 244670+24
- Suicide and self-harm multi-turn appropriate response rate (Claude.ai)Haiku 5.5 by 197190+19
- Election integrity multi-turn (appropriate response rate) - Claude.aiHaiku 5.5 by 118798+11
- SWE-bench MultimodalHaiku 5.5 by 10.919.830.7+10.9
- Election integrity multi-turn (appropriate response rate) - API, without a system promptHaiku 5.5 by 69399+6
- Child safety multi-turn appropriate response rate (API, without a system prompt)Haiku 5.5 by 19596+1
- Child safety multi-turn appropriate response rate (Claude.ai)Haiku 5.5 by 19899+1
- Child safety single-turn harmful prompts decline rate (API)tie99.999.3tie
- Election integrity single-turn benign requests (refusal rate) - API, without a system prompttie00.3tie
- Election integrity single-turn benign requests (refusal rate) - Claude.aitie00.3tie
- Election integrity single-turn harmful requests (harmless rate) - API, without a system prompttie99.8100tie
- Election integrity single-turn harmful requests (harmless rate) - Claude.aitie99.8100tie
- Suicide and self-harm single-turn benign requests refusal rate (Claude.ai)tie0.60.4tie
- Suicide and self-harm single-turn requests posing potential risk harmless rate (Claude.ai)tie99.8100tie
- Suicide and self-harm single-turn requests posing potential risk harmless rate (API, without a system prompt)tie99.599.6tie
- Suicide and self-harm single-turn benign requests refusal rate (API, without a system prompt)tie00tie
Questions people ask
Which is better, Claude Haiku 4.5 or Haiku 5.5?
Haiku 5.5 wins all two areas where both have results: coding and reasoning. Claude Haiku 4.5 wins none.
Which is better for coding?
Haiku 5.5. It wins 2 of the 2 coding tests both models report; Claude Haiku 4.5 wins none.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 21 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.