The two are evenly matched.
Scores updated · 38 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude 3.5 Sonnet pulls ahead
- Common-sense trick questionsSimpleBench+9.4points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+7.4points ahead
Where o1-mini pulls ahead
- Recent programming contest problemsLiveCodeBench v6+9.6points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4.3points ahead
Every test, side by side
All 38 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- LiveCodeBench v6o1-mini by 9.637.246.8+9.6
- SWE-bench VerifiedClaude 3.5 Sonnet by 7.44941.6+7.4
- SciCodetie31.632.3tie
ReasoningEven
- SimpleBenchClaude 3.5 Sonnet by 9.427.518.1+9.4
- GPQA Diamondo1-mini by 4.35660.3+4.3
- Humanity's Last Examtie3.23.6tie
Other results32 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforceso1-mini by 1103 rating points7171820+1103 rating
- Codeforces (Rating)o1-mini by 1103 rating points7171820+1103 rating
- Codeforces (Percentile)o1-mini by 73.120.393.4+73.1
- CNMO 2024o1-mini by 54.513.167.6+54.5
- AIME 2024o1-mini by 47.61663.6+47.6
- AIME 2025o1-mini by 47.43.350.7+47.4
- AIR-Bench 2024Claude 3.5 Sonnet by 40.685.945.3+40.6
- SimpleQAClaude 3.5 Sonnet by 21.428.47+21.4
- LiveCodeBench (Pass@1-COT)o1-mini by 2033.853.8+20
- MATHo1-mini by 18.971.190+18.9
- Chinese SimpleQA (C-SimpleQA)Claude 3.5 Sonnet by 15.155.440.3+15.1
- Aider-PolyglotClaude 3.5 Sonnet by 12.445.332.9+12.4
- MATH-500 (EM)o1-mini by 11.778.390+11.7
- MATH500 (Pass@1)o1-mini by 11.778.390+11.7
- HarmBenchClaude 3.5 Sonnet by 9.698.188.5+9.6
- C-EvalClaude 3.5 Sonnet by 7.876.768.9+7.8
- Arena-Hard (GPT-4-1106 judge)o1-mini by 6.885.292+6.8
- AlpacaEval2.0 (LC-winrate)o1-mini by 5.85257.8+5.8
- CLUEWSCo1-mini by 4.585.489.9+4.5
- DROP (3-shot F1)Claude 3.5 Sonnet by 4.488.383.9+4.4
- FRAMES (Acc.)o1-mini by 4.472.576.9+4.4
- MMLUClaude 3.5 Sonnet by 3.588.785.2+3.5
- simple_safety_testsClaude 3.5 Sonnet by 310097+3
- SuperGPQAClaude 3.5 Sonnet by 348.245.2+3
- MMLU-Proo1-mini by 2.777.680.3+2.7
- AA Intelligenceo1-mini by 2.67.29.8+2.6
- MMLU-ReduxClaude 3.5 Sonnet by 2.288.986.7+2.2
- bbqo1-mini by 294.996.9+2
- IFEvalClaude 3.5 Sonnet by 1.786.584.8+1.7
- anthropic_red_teamClaude 3.5 Sonnet by 1.599.898.3+1.5
- XSTesto1-mini by 1.495.697+1.4
- HumanEvalClaude 3.5 Sonnet by 1.393.792.4+1.3
Questions people ask
Which is better, Claude 3.5 Sonnet or o1-mini?
Neither wins more tests than the other in any of the two areas where both have results.
Which is better for coding?
Neither. They win 1 coding test each of the 3 both models report, and 1 is a tie.
How do you compare the two?
We use the 38 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 32 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.