o1-mini is the stronger all-rounder.
Scores updated · 38 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where o1-mini pulls ahead
- Recent programming contest problemsLiveCodeBench v6+15.9points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+8.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+7.7points ahead
Where GPT-4o pulls ahead
- Code for real scientific research problemsSciCode+1points ahead
Every test, side by side
All 38 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Codingo1-mini
- LiveCodeBench v6o1-mini by 15.930.946.8+15.9
- SWE-bench Verifiedo1-mini by 8.433.241.6+8.4
- SciCodeGPT-4o by 133.332.3+1
Reasoningo1-mini
- GPQA Diamondo1-mini by 7.752.660.3+7.7
- Humanity's Last Examo1-mini by 1.81.83.6+1.8
Other results33 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforceso1-mini by 1061 rating points7591820+1061 rating
- Codeforces (Rating)o1-mini by 1061 rating points7591820+1061 rating
- Codeforces (Percentile)o1-mini by 69.823.693.4+69.8
- CNMO 2024o1-mini by 56.810.867.6+56.8
- AIME 2024o1-mini by 54.39.363.6+54.3
- AIME 2025o1-mini by 3911.750.7+39
- SimpleQAGPT-4o by 31.238.27+31.2
- LiveCodeBench (Pass@1-COT)o1-mini by 19.634.253.8+19.6
- Chinese SimpleQA (C-SimpleQA)GPT-4o by 18.458.740.3+18.4
- AIR-Bench 2024GPT-4o by 17.162.445.3+17.1
- MATH-500 (EM)o1-mini by 15.474.690+15.4
- MATH500 (Pass@1)o1-mini by 15.474.690+15.4
- Arena-Hard (GPT-4-1106 judge)o1-mini by 11.680.492+11.6
- MMLUo1-mini by 11.473.885.2+11.4
- ARC-AGI-1o1-mini by 9.54.514+9.5
- C-EvalGPT-4o by 7.17668.9+7.1
- AlpacaEval2.0 (LC-winrate)o1-mini by 6.751.157.8+6.7
- HarmBencho1-mini by 5.682.988.5+5.6
- MMLU-Proo1-mini by 5.674.780.3+5.6
- MATHo1-mini by 4.785.390+4.7
- IFEvalo1-mini by 3.88184.8+3.8
- FRAMES (Acc.)GPT-4o by 3.680.576.9+3.6
- SuperGPQAo1-mini by 2.842.445.2+2.8
- AA Intelligenceo1-mini by 2.57.39.8+2.5
- Aider-Polygloto1-mini by 2.230.732.9+2.2
- HumanEvalo1-mini by 2.290.292.4+2.2
- CLUEWSCo1-mini by 287.989.9+2
- bbqo1-mini by 1.895.196.9+1.8
- simple_safety_testsGPT-4o by 1.598.597+1.5
- MMLU-ReduxGPT-4o by 1.38886.7+1.3
- anthropic_red_teamtie99.198.3tie
- XSTesttie97.397tie
- DROP (3-shot F1)tie83.783.9tie
Questions people ask
Which is better, GPT-4o or o1-mini?
o1-mini wins all two areas where both have results: coding and reasoning. GPT-4o wins none.
Which is better for coding?
o1-mini. It wins 2 of the 3 coding tests both models report; GPT-4o wins 1.
How do you compare the two?
We use the 38 benchmark tests both models have published scores on. The verdict counts the 5 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 33 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.