o1-mini is the stronger all-rounder.
Scores updated · 26 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where o1-mini pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+17.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+11.2points ahead
- Code for real scientific research problemsSciCode+5.6points ahead
Where Qwen2.5 Instruct 72B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 26 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Codingo1-mini
- SWE-bench Verifiedo1-mini by 17.841.623.8+17.8
- SciCodeo1-mini by 5.632.326.7+5.6
Reasoningo1-mini
- GPQA Diamondo1-mini by 11.260.349.1+11.2
- Humanity's Last Examtie3.63.6tie
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforceso1-mini by 1795 rating points182024.8+1795 rating
- AIME 2024o1-mini by 40.363.623.3+40.3
- C-EvalQwen2.5 Instruct 72B by 20.368.989.2+20.3
- HarmBencho1-mini by 15.788.572.8+15.7
- AIR-Bench 2024Qwen2.5 Instruct 72B by 13.745.359+13.7
- MATH-500 (EM)o1-mini by 109080+10
- MMLU-Proo1-mini by 9.280.371.1+9.2
- MMLUo1-mini by 8.285.277+8.2
- Chinese SimpleQA (C-SimpleQA)Qwen2.5 Instruct 72B by 8.140.348.4+8.1
- CLUEWSCo1-mini by 7.489.982.5+7.4
- DROP (3-shot F1)o1-mini by 7.283.976.7+7.2
- FRAMES (Acc.)o1-mini by 7.176.969.8+7.1
- HumanEvalo1-mini by 5.892.486.6+5.8
- simple_safety_testsQwen2.5 Instruct 72B by 397100+3
- AA Intelligenceo1-mini by 2.19.87.7+2.1
- SimpleQAQwen2.5 Instruct 72B by 2.179.1+2.1
- MATHo1-mini by 1.69088.4+1.6
- bbqo1-mini by 1.596.995.4+1.5
- anthropic_red_teamQwen2.5 Instruct 72B by 1.398.399.6+1.3
- XSTesttie9797.9tie
- IFEvaltie84.884.1tie
- MMLU-Reduxtie86.786.8tie
Questions people ask
Which is better, o1-mini or Qwen2.5 Instruct 72B?
o1-mini wins all two areas where both have results: coding and reasoning. Qwen2.5 Instruct 72B wins none.
Which is better for coding?
o1-mini. It wins 2 of the 2 coding tests both models report; Qwen2.5 Instruct 72B wins none.
How do you compare the two?
We use the 26 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.