DeepSeek-V3 is the stronger all-rounder.
Scores updated · 62 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning and long documents rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3 pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+30.2points ahead
- Questions about very long textsLongBench v2+17.1points ahead
Where DeepSeek-V2 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 62 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results60 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- CodeforcesDeepSeek-V3 by 1117 rating points17.51134+1117 rating
- MATHDeepSeek-V3 by 46.843.490.2+46.8
- MultiPL-EDeepSeek-V3 by 38.744.483.1+38.7
- LiveCodeBenchDeepSeek-V3 by 37.611.649.2+37.6
- AIME 2024DeepSeek-V3 by 34.64.639.2+34.6
- MATH-500 (EM)DeepSeek-V3 by 33.956.390.2+33.9
- IFEvalDeepSeek-V3 by 28.457.786.1+28.4
- MMLU-ProDeepSeek-V3 by 24.551.475.9+24.5
- AGIEvalDeepSeek-V3 by 22.157.579.6+22.1
- HumanEvalDeepSeek-V3 by 21.943.365.2+21.9
- CRUXEval-O (output prediction)DeepSeek-V3 by 2049.869.8+20
- Chinese SimpleQA (C-SimpleQA)DeepSeek-V3 by 19.548.568+19.5
- Aider-Edit (Acc.)DeepSeek-V3 by 19.460.379.7+19.4
- MATH (EM)DeepSeek-V3 by 18.243.461.6+18.2
- MGSMDeepSeek-V3 by 16.263.679.8+16.2
- MGSM (EM)DeepSeek-V3 by 16.263.679.8+16.2
- MMMLUDeepSeek-V3 by 15.46479.4+15.4
- MMMLU-non-English (Acc.)DeepSeek-V3 by 15.46479.4+15.4
- CRUXEval-I (input prediction)DeepSeek-V3 by 14.852.567.3+14.8
- GSM8KDeepSeek-V3 by 14.879.294+14.8
- HumanEval-Mul (Pass@1)DeepSeek-V3 by 13.369.382.6+13.3
- MMLU-Pro (Acc.)DeepSeek-V3 by 1351.464.4+13
- MMLUDeepSeek-V3 by 12.975.688.5+12.9
- CMath (EM)DeepSeek-V3 by 1278.790.7+12
- DROPDeepSeek-V3 by 11.580.191.6+11.5
- MMLU-ReduxDeepSeek-V3 by 11.277.989.1+11.2
- MMLU-Redux (Acc.)DeepSeek-V3 by 10.675.686.2+10.6
- MBPPDeepSeek-V3 by 10.46575.4+10.4
- CLUEWSCDeepSeek-V3 by 8.98290.9+8.9
- BBHDeepSeek-V3 by 8.778.887.5+8.7
- BBH (EM)DeepSeek-V3 by 8.778.887.5+8.7
- MMLU (Acc.)DeepSeek-V3 by 8.778.487.1+8.7
- DROP (3-shot F1)DeepSeek-V3 by 8.68391.6+8.6
- DROP (F1)DeepSeek-V3 by 8.680.489+8.6
- LiveCodeBench-Base (Pass@1)DeepSeek-V3 by 7.811.619.4+7.8
- GSM8K (EM)DeepSeek-V3 by 7.781.689.3+7.7
- FRAMES (Acc.)DeepSeek-V3 by 6.466.973.3+6.4
- RACE-MiddleDeepSeek-V2 by 673.167.1+6
- C-EvalDeepSeek-V3 by 5.181.486.5+5.1
- CMMLUDeepSeek-V3 by 4.88488.8+4.8
- CMMLU (Acc.)DeepSeek-V3 by 4.88488.8+4.8
- ARC-ChallengeDeepSeek-V3 by 3.192.295.3+3.1
- ARC-Challenge (Acc.)DeepSeek-V3 by 3.192.295.3+3.1
- TriviaQA (EM)DeepSeek-V3 by 2.98082.9+2.9
- HellaSwagDeepSeek-V3 by 1.887.188.9+1.8
- HellaSwag (Acc.)DeepSeek-V3 by 1.887.188.9+1.8
- NaturalQuestionsDeepSeek-V3 by 1.438.640+1.4
- NaturalQuestions (EM)DeepSeek-V3 by 1.438.640+1.4
- WinoGrandeDeepSeek-V2 by 1.486.384.9+1.4
- WinoGrande (Acc.)DeepSeek-V2 by 1.486.384.9+1.4
- ARC-Easy (Acc.)DeepSeek-V3 by 1.397.698.9+1.3
- RACE-HighDeepSeek-V2 by 1.352.651.3+1.3
- C3DeepSeek-V3 by 1.277.478.6+1.2
- C3 (Acc.)DeepSeek-V3 by 1.277.478.6+1.2
- CMRCDeepSeek-V2 by 1.177.476.3+1.1
- CCPMDeepSeek-V2 by 19392+1
- PIQAtie83.984.7tie
- PIQA (Acc.)tie83.984.7tie
- Pile-test (BPB)lower is bettertie0.60.5tie
- The Pile (Test, BPB)lower is bettertie0.60.5tie
Questions people ask
Which is better, DeepSeek-V2 or DeepSeek-V3?
DeepSeek-V3 wins all two areas where both have results: reasoning and long documents. DeepSeek-V2 wins none.
How do you compare the two?
We use the 62 benchmark tests both models have published scores on. The verdict counts the 2 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 60 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.