Sonnet 5.5 is the stronger all-rounder.
Scores updated · 28 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Sonnet 5.5 pulls ahead
- Long, multi-file coding tasks in real codebasesSWE-bench Pro+16.5points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+7.1points ahead
- Fixes real GitHub issues in many programming languagesSWE-bench Multilingual+6.6points ahead
Where Haiku 5.5 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 28 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingSonnet 5.5
- SWE-bench ProSonnet 5.5 by 16.564.881.3+16.5
- SWE-bench MultilingualSonnet 5.5 by 6.683.790.3+6.6
Other results25 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Terminal-Bench-Science 0.1Sonnet 5.5 by 39.320.659.9+39.3
- Terminal-Bench 4.0Sonnet 5.5 by 31.439.270.6+31.4
- SWE-bench MultimodalSonnet 5.5 by 23.630.754.3+23.6
- GDPval-AA 2.1Sonnet 5.5 by 220 rating points16201840+220 rating
- Election integrity multi-turn (appropriate response rate) - API, without a system promptHaiku 5.5 by 139986+13
- Election integrity multi-turn (appropriate response rate) - Claude.aiHaiku 5.5 by 129886+12
- OSWorld 2.1 (offline subset)Sonnet 5.5 by 11.97283.9+11.9
- OSWorld 2.1Sonnet 5.5 by 11.572.483.9+11.5
- Child safety multi-turn appropriate response rate (API, without a system prompt)Haiku 5.5 by 109686+10
- Suicide and self-harm multi-turn appropriate response rate (API, without a system prompt)Haiku 5.5 by 107060+10
- Suicide and self-harm multi-turn appropriate response rate (Claude.ai)Sonnet 5.5 by 1090100+10
- FrontierCode 1.1 (Main) xhighSonnet 5.5 by 6.345.852.1+6.3
- HealthBench Professional length-adjustedSonnet 5.5 by 4.464.869.2+4.4
- Program BenchHaiku 5.5 by 2.38279.7+2.3
- Child safety multi-turn appropriate response rate (Claude.ai)Haiku 5.5 by 19998+1
- FrontierCode Extendedtie58.459.1tie
- Suicide and self-harm single-turn requests posing potential risk harmless rate (API, without a system prompt)tie99.699.1tie
- Election integrity single-turn benign requests (refusal rate) - API, without a system prompttie0.30tie
- Election integrity single-turn benign requests (refusal rate) - Claude.aitie0.30tie
- FrontierCode 1.1 (Main) maxtie46.446.2tie
- Suicide and self-harm single-turn requests posing potential risk harmless rate (Claude.ai)tie10099.8tie
- Election integrity single-turn harmful requests (harmless rate) - API, without a system prompttie100100tie
- Election integrity single-turn harmful requests (harmless rate) - Claude.aitie100100tie
- Suicide and self-harm single-turn benign requests refusal rate (API, without a system prompt)tie00tie
- Suicide and self-harm single-turn benign requests refusal rate (Claude.ai)tie0.40.4tie
Questions people ask
Which is better, Haiku 5.5 or Sonnet 5.5?
Sonnet 5.5 wins all two areas where both have results: coding and reasoning. Haiku 5.5 wins none.
Which is better for coding?
Sonnet 5.5. It wins 2 of the 2 coding tests both models report; Haiku 5.5 wins none.
How do you compare the two?
We use the 28 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 25 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.