o1-mini wins more areas, narrowly.
Scores updated · 27 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where o1-mini pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+17.1points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+8.8points ahead
- Code for real scientific research problemsSciCode+2.4points ahead
Where Llama 3.1 Instruct 405B pulls ahead
- Common-sense trick questionsSimpleBench+4.9points ahead
Every test, side by side
All 27 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Codingo1-mini
- SWE-bench Verifiedo1-mini by 17.124.541.6+17.1
- SciCodeo1-mini by 2.429.932.3+2.4
ReasoningEven
- GPQA Diamondo1-mini by 8.851.560.3+8.8
- SimpleBenchLlama 3.1 Instruct 405B by 4.92318.1+4.9
- Humanity's Last Examtie43.6tie
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforceso1-mini by 1795 rating points25.31820+1795 rating
- AIME 2024o1-mini by 40.323.363.6+40.3
- HarmBencho1-mini by 25.862.788.5+25.8
- MATHo1-mini by 16.273.890+16.2
- MATH-500 (EM)o1-mini by 16.273.890+16.2
- AIR-Bench 2024Llama 3.1 Instruct 405B by 13.358.645.3+13.3
- Chinese SimpleQA (C-SimpleQA)Llama 3.1 Instruct 405B by 10.150.440.3+10.1
- SimpleQALlama 3.1 Instruct 405B by 10.117.17+10.1
- MMLU-Proo1-mini by 773.380.3+7
- CLUEWSCo1-mini by 6.98389.9+6.9
- FRAMES (Acc.)o1-mini by 6.97076.9+6.9
- DROP (3-shot F1)Llama 3.1 Instruct 405B by 4.888.783.9+4.8
- IFEvalLlama 3.1 Instruct 405B by 3.888.684.8+3.8
- C-EvalLlama 3.1 Instruct 405B by 3.672.568.9+3.6
- HumanEvalo1-mini by 3.48992.4+3.4
- MMLULlama 3.1 Instruct 405B by 3.488.685.2+3.4
- AA Intelligenceo1-mini by 2.57.39.8+2.5
- bbqo1-mini by 2.494.596.9+2.4
- anthropic_red_teamo1-mini by 1.896.598.3+1.8
- simple_safety_testsLlama 3.1 Instruct 405B by 1.898.897+1.8
- XSTesto1-mini by 1.195.997+1.1
- MMLU-Reduxtie86.286.7tie
Questions people ask
Which is better, Llama 3.1 Instruct 405B or o1-mini?
o1-mini wins one of the two areas where both have results: coding. Llama 3.1 Instruct 405B wins none. They are level on reasoning.
Which is better for coding?
o1-mini. It wins 2 of the 2 coding tests both models report; Llama 3.1 Instruct 405B wins none.
How do you compare the two?
We use the 27 benchmark tests both models have published scores on. The verdict counts the 5 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.