Claude 3 Opus wins more areas, narrowly.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude 3 Opus pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+11.8points ahead
Where Qwen2 Instruct (72B) pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningClaude 3 Opus
- GPQA DiamondClaude 3 Opus by 11.848.937.1+11.8
- Humanity's Last Examtie2.83.7tie
Other results18 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- NarrativeQAQwen2 Instruct (72B) by 37.635.172.7+37.6
- ARC-ChallengeClaude 3 Opus by 27.596.468.9+27.5
- AIR-Bench 2024Claude 3 Opus by 22.384.462.1+22.3
- HarmBenchClaude 3 Opus by 20.697.476.8+20.6
- MATHQwen2 Instruct (72B) by 18.960.179+18.9
- MMLUClaude 3 Opus by 9.986.876.9+9.9
- HellaSwagClaude 3 Opus by 7.895.487.6+7.8
- naturalquestions_closedbookClaude 3 Opus by 5.14439+5.1
- BBHClaude 3 Opus by 4.486.882.4+4.4
- XSTestQwen2 Instruct (72B) by 4.492.596.9+4.4
- MMLU-ProClaude 3 Opus by 4.168.564.4+4.1
- GSM8KClaude 3 Opus by 3.99591.1+3.9
- AA IntelligenceClaude 3 Opus by 2.48.76.3+2.4
- simple_safety_testsClaude 3 Opus by 1.510098.5+1.5
- bbqQwen2 Instruct (72B) by 1.19495.1+1.1
- HumanEvalQwen2 Instruct (72B) by 1.184.986+1.1
- anthropic_red_teamtie99.899.1tie
- OpenBookQAtie95.695.4tie
Questions people ask
Which is better, Claude 3 Opus or Qwen2 Instruct (72B)?
Claude 3 Opus wins one of the two areas where both have results: reasoning. Qwen2 Instruct (72B) wins none. They are level on coding.
Which is better for coding?
Neither. All 1 coding tests both models report are ties.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 18 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.