DeepSeek-V3.1-Base wins more areas, narrowly.
Scores updated · 41 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and reasoning rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3.1-Base pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+19.2points ahead
Where DeepSeek-V3.2-Exp-Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 41 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results39 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- NIAH-Multi (32K)DeepSeek-V3.1-Base by 14.199.785.6+14.1
- NIAH-Multi (64K)DeepSeek-V3.1-Base by 12.798.685.9+12.7
- HumanEvalDeepSeek-V3.1-Base by 10.672.561.9+10.6
- GSM-Infinite Hard (16K)DeepSeek-V3.2-Exp-Base by 8.941.550.4+8.9
- GSM8KDeepSeek-V3.1-Base by 791.484.4+7
- GSM-Infinite Hard (32K)DeepSeek-V3.2-Exp-Base by 6.438.845.2+6.4
- GPQA (unspecified)DeepSeek-V3.1-Base by 5.843.137.3+5.8
- MMLU-ProDeepSeek-V3.2-Exp-Base by 4.558.863.3+4.5
- aime_2024_2025_combinedDeepSeek-V3.2-Exp-Base by 3.221.624.8+3.2
- HumanEval+DeepSeek-V3.2-Exp-Base by 3.164.667.7+3.1
- GSM-Infinite Hard (Tools-allowed)DeepSeek-V3.1-Base by 328.725.7+3
- GSM-Infinite Hard (128K)DeepSeek-V3.1-Base by 328.725.7+3
- Chinese SimpleQA (C-SimpleQA)DeepSeek-V3.1-Base by 2.970.968+2.9
- NIAH-Multi (Tools-allowed)DeepSeek-V3.1-Base by 2.997.294.3+2.9
- NIAH-Multi (128K)DeepSeek-V3.1-Base by 2.997.294.3+2.9
- MATHDeepSeek-V3.1-Base by 2.562.660.1+2.5
- WinoGrandeDeepSeek-V3.1-Base by 2.585.983.4+2.5
- MBPP+DeepSeek-V3.1-Base by 2.472.269.8+2.4
- MBPP+ (EvalPlus-augmented)DeepSeek-V3.1-Base by 2.472.269.8+2.4
- GSM-Infinite Hard (64K)DeepSeek-V3.1-Base by 2.134.732.6+2.1
- MultiPL-E MBPPDeepSeek-V3.1-Base by 1.952.550.6+1.9
- CRUXEval-I (input prediction)DeepSeek-V3.2-Exp-Base by 1.862.163.9+1.8
- CRUXEval-O (output prediction)DeepSeek-V3.1-Base by 1.576.474.9+1.5
- SuperGPQADeepSeek-V3.2-Exp-Base by 1.342.343.6+1.3
- C-EvalDeepSeek-V3.2-Exp-Base by 19091+1
- MBPPDeepSeek-V3.2-Exp-Base by 174.675.6+1
- SimpleQAtie26.327tie
- BBHtie88.288.7tie
- ARC-Challengetie95.695.2tie
- MMLUtie87.487.8tie
- MMLU-Reduxtie9090.4tie
- TriviaQAtie83.583.9tie
- DROPtie86.386.6tie
- HellaSwagtie89.289.4tie
- MultiPL-E HumanEvaltie45.945.7tie
- BigCodeBenchtie6362.9tie
- CMMLUtie88.888.9tie
- GlobalMMLUtie81.982tie
- Includetie77.277.2tie
Questions people ask
Which is better, DeepSeek-V3.1-Base or DeepSeek-V3.2-Exp-Base?
DeepSeek-V3.1-Base wins one of the two areas where both have results: reasoning. DeepSeek-V3.2-Exp-Base wins none. They are level on coding.
Which is better for coding?
Neither. All 1 coding tests both models report are ties.
How do you compare the two?
We use the 41 benchmark tests both models have published scores on. The verdict counts the 2 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 39 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.