DeepSeek-V3 is the stronger all-rounder.
Scores updated · 22 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and reasoning rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3 pulls ahead
- Recent programming contest problemsLiveCodeBench v6+22.1points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+14.5points ahead
Where DeepSeek-V3.1-Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 22 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results20 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MATHDeepSeek-V3 by 27.690.262.6+27.6
- MMLU-ProDeepSeek-V3 by 17.175.958.8+17.1
- SuperGPQADeepSeek-V3 by 11.453.742.3+11.4
- HumanEvalDeepSeek-V3.1-Base by 7.365.272.5+7.3
- CRUXEval-O (output prediction)DeepSeek-V3.1-Base by 6.669.876.4+6.6
- MBPP+ (EvalPlus-augmented)DeepSeek-V3 by 6.678.872.2+6.6
- DROPDeepSeek-V3 by 5.391.686.3+5.3
- CRUXEval-I (input prediction)DeepSeek-V3 by 5.267.362.1+5.2
- C-EvalDeepSeek-V3.1-Base by 3.586.590+3.5
- Chinese SimpleQA (C-SimpleQA)DeepSeek-V3.1-Base by 2.96870.9+2.9
- GSM8KDeepSeek-V3 by 2.69491.4+2.6
- SimpleQADeepSeek-V3.1-Base by 1.424.926.3+1.4
- MMLUDeepSeek-V3 by 1.188.587.4+1.1
- WinoGrandeDeepSeek-V3.1-Base by 184.985.9+1
- MMLU-Reduxtie89.190tie
- MBPPtie75.474.6tie
- BBHtie87.588.2tie
- ARC-Challengetie95.395.6tie
- HellaSwagtie88.989.2tie
- CMMLUtie88.888.8tie
Questions people ask
Which is better, DeepSeek-V3 or DeepSeek-V3.1-Base?
DeepSeek-V3 wins all two areas where both have results: coding and reasoning. DeepSeek-V3.1-Base wins none.
Which is better for coding?
DeepSeek-V3. It wins the one coding test both models report.
How do you compare the two?
We use the 22 benchmark tests both models have published scores on. The verdict counts the 2 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 20 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.