Qwen2 72B is the stronger all-rounder.
Scores updated · 17 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen2 72B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+8.3points ahead
Where Llama 3 Instruct 8B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results16 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- C-EvalQwen2 72B by 41.549.591+41.5
- Multi-MathematicsQwen2 72B by 39.736.376+39.7
- GSM8KQwen2 72B by 39.649.989.5+39.6
- CMMLUQwen2 72B by 39.350.890.1+39.3
- MultiPL-EQwen2 72B by 3722.659.6+37
- EvalPlusQwen2 72B by 25.140.365.4+25.1
- BBHQwen2 72B by 24.757.782.4+24.7
- MMLUQwen2 72B by 2460.184.2+24
- Theorem QAQwen2 72B by 2122.143.1+21
- MMLU-ProQwen2 72B by 20.235.455.6+20.2
- HellaSwagQwen2 72B by 16.571.187.6+16.5
- ARC-ChallengeLlama 3 Instruct 8B by 13.982.868.9+13.9
- MATHQwen2 72B by 1239.151.1+12
- MBPPQwen2 72B by 9.267.776.9+9.2
- TruthfulQALlama 3 Instruct 8B by 8.463.254.8+8.4
- HumanEvalQwen2 72B by 4.260.464.6+4.2
Questions people ask
Which is better, Llama 3 Instruct 8B or Qwen2 72B?
Qwen2 72B wins the one area where both have results: reasoning. Llama 3 Instruct 8B wins none.
How do you compare the two?
We use the 17 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 16 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.