Qwen2.5 Instruct 72B is the stronger all-rounder.
Scores updated · 62 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning and long documents rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen2.5 Instruct 72B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+13.8points ahead
- Questions about very long textsLongBench v2+7.8points ahead
Where DeepSeek-V2 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 62 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningQwen2.5 Instruct 72B
- GPQA DiamondQwen2.5 Instruct 72B by 13.835.349.1+13.8
Long documentsQwen2.5 Instruct 72B
- LongBench v2Qwen2.5 Instruct 72B by 7.831.639.4+7.8
Other results60 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MATHQwen2.5 Instruct 72B by 4543.488.4+45
- LiveCodeBenchQwen2.5 Instruct 72B by 43.911.655.5+43.9
- HumanEvalQwen2.5 Instruct 72B by 43.343.386.6+43.3
- MultiPL-EQwen2.5 Instruct 72B by 30.744.475.1+30.7
- IFEvalQwen2.5 Instruct 72B by 26.457.784.1+26.4
- MATH-500 (EM)Qwen2.5 Instruct 72B by 23.756.380+23.7
- MBPPQwen2.5 Instruct 72B by 23.26588.2+23.2
- MMLU-ProQwen2.5 Instruct 72B by 19.751.471.1+19.7
- AIME 2024Qwen2.5 Instruct 72B by 18.74.623.3+18.7
- AGIEvalQwen2.5 Instruct 72B by 18.357.575.8+18.3
- GSM8KQwen2.5 Instruct 72B by 16.679.295.8+16.6
- MGSMQwen2.5 Instruct 72B by 12.663.676.2+12.6
- MGSM (EM)Qwen2.5 Instruct 72B by 12.663.676.2+12.6
- MATH (EM)Qwen2.5 Instruct 72B by 1143.454.4+11
- MMMLUQwen2.5 Instruct 72B by 10.86474.8+10.8
- MMMLU-non-English (Acc.)Qwen2.5 Instruct 72B by 10.86474.8+10.8
- CRUXEval-O (output prediction)Qwen2.5 Instruct 72B by 10.149.859.9+10.1
- MMLU-ReduxQwen2.5 Instruct 72B by 8.977.986.8+8.9
- TriviaQA (EM)DeepSeek-V2 by 8.18071.9+8.1
- HumanEval-Mul (Pass@1)Qwen2.5 Instruct 72B by 869.377.3+8
- C-EvalQwen2.5 Instruct 72B by 7.881.489.2+7.8
- MMLU-Redux (Acc.)Qwen2.5 Instruct 72B by 7.675.683.2+7.6
- CodeforcesQwen2.5 Instruct 72B by 7.317.524.8+7.3
- MMLU-Pro (Acc.)Qwen2.5 Instruct 72B by 6.951.458.3+6.9
- GSM8K (EM)Qwen2.5 Instruct 72B by 6.781.688.3+6.7
- CRUXEval-I (input prediction)Qwen2.5 Instruct 72B by 6.652.559.1+6.6
- MMLU (Acc.)Qwen2.5 Instruct 72B by 6.678.485+6.6
- DROP (3-shot F1)DeepSeek-V2 by 6.38376.7+6.3
- CMath (EM)Qwen2.5 Instruct 72B by 5.878.784.5+5.8
- CMMLUQwen2.5 Instruct 72B by 5.58489.5+5.5
- CMMLU (Acc.)Qwen2.5 Instruct 72B by 5.58489.5+5.5
- NaturalQuestionsDeepSeek-V2 by 5.438.633.2+5.4
- NaturalQuestions (EM)DeepSeek-V2 by 5.438.633.2+5.4
- Aider-Edit (Acc.)Qwen2.5 Instruct 72B by 5.160.365.4+5.1
- RACE-MiddleDeepSeek-V2 by 573.168.1+5
- CCPMDeepSeek-V2 by 4.59388.5+4.5
- WinoGrandeDeepSeek-V2 by 486.382.3+4
- WinoGrande (Acc.)DeepSeek-V2 by 486.382.3+4
- DROPDeepSeek-V2 by 3.480.176.7+3.4
- FRAMES (Acc.)Qwen2.5 Instruct 72B by 2.966.969.8+2.9
- ARC-ChallengeQwen2.5 Instruct 72B by 2.392.294.5+2.3
- ARC-Challenge (Acc.)Qwen2.5 Instruct 72B by 2.392.294.5+2.3
- HellaSwagDeepSeek-V2 by 2.387.184.8+2.3
- HellaSwag (Acc.)DeepSeek-V2 by 2.387.184.8+2.3
- RACE-HighDeepSeek-V2 by 2.352.650.3+2.3
- CMRCDeepSeek-V2 by 1.677.475.8+1.6
- MMLUQwen2.5 Instruct 72B by 1.475.677+1.4
- LiveCodeBench-Base (Pass@1)Qwen2.5 Instruct 72B by 1.311.612.9+1.3
- PIQADeepSeek-V2 by 1.383.982.6+1.3
- PIQA (Acc.)DeepSeek-V2 by 1.383.982.6+1.3
- BBHQwen2.5 Instruct 72B by 178.879.8+1
- BBH (EM)Qwen2.5 Instruct 72B by 178.879.8+1
- ARC-Easy (Acc.)tie97.698.4tie
- C3tie77.476.7tie
- C3 (Acc.)tie77.476.7tie
- CLUEWSCtie8282.5tie
- DROP (F1)tie80.480.6tie
- Chinese SimpleQA (C-SimpleQA)tie48.548.4tie
- Pile-test (BPB)lower is bettertie0.60.6tie
- The Pile (Test, BPB)lower is bettertie0.60.6tie
Questions people ask
Which is better, DeepSeek-V2 or Qwen2.5 Instruct 72B?
Qwen2.5 Instruct 72B wins all two areas where both have results: reasoning and long documents. DeepSeek-V2 wins none.
How do you compare the two?
We use the 62 benchmark tests both models have published scores on. The verdict counts the 2 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 60 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.