Qwen2.5 Instruct 72B is the stronger all-rounder.
Scores updated · 23 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen2.5 Instruct 72B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+17.3points ahead
Where DeepSeek-V3.2-Exp-Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 23 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningQwen2.5 Instruct 72B
- GPQA DiamondQwen2.5 Instruct 72B by 17.331.849.1+17.3
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- OpenBookQAQwen2.5 Instruct 72B by 4848.296.2+48
- MATHQwen2.5 Instruct 72B by 28.360.188.4+28.3
- HumanEvalQwen2.5 Instruct 72B by 24.761.986.6+24.7
- Chinese SimpleQA (C-SimpleQA)DeepSeek-V3.2-Exp-Base by 19.66848.4+19.6
- SimpleQADeepSeek-V3.2-Exp-Base by 17.9279.1+17.9
- CRUXEval-O (output prediction)DeepSeek-V3.2-Exp-Base by 1574.959.9+15
- MBPPQwen2.5 Instruct 72B by 12.675.688.2+12.6
- GSM8KQwen2.5 Instruct 72B by 11.484.495.8+11.4
- MMLUDeepSeek-V3.2-Exp-Base by 10.887.877+10.8
- DROPDeepSeek-V3.2-Exp-Base by 9.986.676.7+9.9
- BBHDeepSeek-V3.2-Exp-Base by 8.988.779.8+8.9
- MMLU-ProQwen2.5 Instruct 72B by 7.863.371.1+7.8
- MBPP+ (EvalPlus-augmented)Qwen2.5 Instruct 72B by 7.269.877+7.2
- MGSMDeepSeek-V3.2-Exp-Base by 6.182.376.2+6.1
- CRUXEval-I (input prediction)DeepSeek-V3.2-Exp-Base by 4.863.959.1+4.8
- HellaSwagDeepSeek-V3.2-Exp-Base by 4.689.484.8+4.6
- MMLU-ReduxDeepSeek-V3.2-Exp-Base by 3.690.486.8+3.6
- PIQADeepSeek-V3.2-Exp-Base by 2.585.182.6+2.5
- C-EvalDeepSeek-V3.2-Exp-Base by 1.89189.2+1.8
- WinoGrandeDeepSeek-V3.2-Exp-Base by 1.183.482.3+1.1
- ARC-Challengetie95.294.5tie
- CMMLUtie88.989.5tie
Questions people ask
Which is better, DeepSeek-V3.2-Exp-Base or Qwen2.5 Instruct 72B?
Qwen2.5 Instruct 72B wins the one area where both have results: reasoning. DeepSeek-V3.2-Exp-Base wins none.
How do you compare the two?
We use the 23 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.