Llama 3.1 Instruct 405B is the stronger all-rounder.
Scores updated · 25 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Llama 3.1 Instruct 405B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+19.7points ahead
Where DeepSeek-V3.2-Exp-Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 25 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningLlama 3.1 Instruct 405B
- GPQA DiamondLlama 3.1 Instruct 405B by 19.731.851.5+19.7
Other results24 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- OpenBookQALlama 3.1 Instruct 405B by 45.848.294+45.8
- MultiPL-E HumanEvalLlama 3.1 Instruct 405B by 29.545.775.2+29.5
- HumanEvalLlama 3.1 Instruct 405B by 27.161.989+27.1
- C-EvalDeepSeek-V3.2-Exp-Base by 18.59172.5+18.5
- Chinese SimpleQA (C-SimpleQA)DeepSeek-V3.2-Exp-Base by 17.66850.4+17.6
- CMMLUDeepSeek-V3.2-Exp-Base by 15.288.973.7+15.2
- MultiPL-E MBPPLlama 3.1 Instruct 405B by 15.150.665.7+15.1
- CRUXEval-O (output prediction)DeepSeek-V3.2-Exp-Base by 1574.959.9+15
- MATHLlama 3.1 Instruct 405B by 13.760.173.8+13.7
- GSM8KLlama 3.1 Instruct 405B by 12.484.496.8+12.4
- MMLU-ProLlama 3.1 Instruct 405B by 1063.373.3+10
- SimpleQADeepSeek-V3.2-Exp-Base by 9.92717.1+9.9
- MGSMLlama 3.1 Instruct 405B by 9.382.391.6+9.3
- MBPPDeepSeek-V3.2-Exp-Base by 7.275.668.4+7.2
- BBHDeepSeek-V3.2-Exp-Base by 5.888.782.9+5.8
- CRUXEval-I (input prediction)DeepSeek-V3.2-Exp-Base by 5.463.958.5+5.4
- MMLU-ReduxDeepSeek-V3.2-Exp-Base by 4.290.486.2+4.2
- MBPP+ (EvalPlus-augmented)Llama 3.1 Instruct 405B by 3.269.873+3.2
- DROPDeepSeek-V3.2-Exp-Base by 1.886.684.8+1.8
- WinoGrandeLlama 3.1 Instruct 405B by 1.883.485.2+1.8
- ARC-ChallengeLlama 3.1 Instruct 405B by 1.795.296.9+1.7
- MMLUtie87.888.6tie
- PIQAtie85.185.9tie
- HellaSwagtie89.489.2tie
Questions people ask
Which is better, DeepSeek-V3.2-Exp-Base or Llama 3.1 Instruct 405B?
Llama 3.1 Instruct 405B wins the one area where both have results: reasoning. DeepSeek-V3.2-Exp-Base wins none.
How do you compare the two?
We use the 25 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 24 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.