Claude Sonnet 4 is the stronger all-rounder.
Scores updated · 25 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude Sonnet 4 pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+31.1points ahead
- Common-sense trick questionsSimpleBench+27.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+17.4points ahead
Where o1-mini pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 25 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Sonnet 4
- SWE-bench VerifiedClaude Sonnet 4 by 31.172.741.6+31.1
- SciCodeClaude Sonnet 4 by 7.74032.3+7.7
- LiveCodeBench v6Claude Sonnet 4 by 1.748.546.8+1.7
ReasoningClaude Sonnet 4
- SimpleBenchClaude Sonnet 4 by 27.445.518.1+27.4
- GPQA DiamondClaude Sonnet 4 by 17.477.760.3+17.4
- Humanity's Last ExamClaude Sonnet 4 by 7.110.73.6+7.1
Other results19 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIR-Bench 2024Claude Sonnet 4 by 4388.345.3+43
- Aider-PolyglotClaude Sonnet 4 by 37.870.732.9+37.8
- ARC-AGI-1Claude Sonnet 4 by 264014+26
- AIME 2024o1-mini by 20.243.463.6+20.2
- AIME 2025Claude Sonnet 4 by 19.870.550.7+19.8
- SuperGPQAClaude Sonnet 4 by 10.555.745.2+10.5
- HarmBenchClaude Sonnet 4 by 9.59888.5+9.5
- AA IntelligenceClaude Sonnet 4 by 9.118.99.8+9.1
- SimpleQAClaude Sonnet 4 by 8.915.97+8.9
- CNMO 2024o1-mini by 7.260.467.6+7.2
- MMLU-ReduxClaude Sonnet 4 by 6.993.686.7+6.9
- MMLUClaude Sonnet 4 by 6.391.585.2+6.3
- MATH-500 (EM)Claude Sonnet 4 by 49490+4
- MMLU-ProClaude Sonnet 4 by 3.483.780.3+3.4
- simple_safety_testsClaude Sonnet 4 by 310097+3
- IFEvalClaude Sonnet 4 by 2.887.684.8+2.8
- anthropic_red_teamtie99.298.3tie
- XSTesttie96.697tie
- bbqtie96.696.9tie
Questions people ask
Which is better, Claude Sonnet 4 or o1-mini?
Claude Sonnet 4 wins all two areas where both have results: coding and reasoning. o1-mini wins none.
Which is better for coding?
Claude Sonnet 4. It wins 3 of the 3 coding tests both models report; o1-mini wins none.
How do you compare the two?
We use the 25 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 19 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.