Claude 3 Opus is the stronger all-rounder.
Scores updated · 38 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and images and charts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude 3 Opus pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+8.9points ahead
- Math problems shown in pictures and chartsMathVista+2.6points ahead
Where Claude 3 Sonnet pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 38 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningClaude 3 Opus
- GPQA DiamondClaude 3 Opus by 8.948.940+8.9
- Humanity's Last Examtie2.83.6tie
Images and chartsClaude 3 Opus
- MathVistaClaude 3 Opus by 2.650.547.9+2.6
Other results34 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- naturalquestions_closedbookClaude 3 Opus by 41.3442.8+41.3
- NarrativeQAClaude 3 Opus by 2435.111.1+24
- MATHClaude 3 Opus by 1760.143.1+17
- GPQA (Diamond) Maj@32 5-shot CoTClaude 3 Opus by 13.259.546.3+13.2
- HumanEvalClaude 3 Opus by 11.984.973+11.9
- HumanEval 0-shotClaude 3 Opus by 11.984.973+11.9
- MMLU-ProClaude 3 Opus by 11.768.556.8+11.7
- BBQ Disambig Accuracy (%)Claude 3 Sonnet by 11.47990.4+11.4
- GPQA (Diamond) 0-shot CoTClaude 3 Opus by 1050.440.4+10
- MMLU 0-shot CoTClaude 3 Opus by 8.685.777.1+8.6
- MMLUClaude 3 Opus by 8.586.878.3+8.5
- MGSMClaude 3 Opus by 7.290.783.5+7.2
- MGSM 0-shot CoTClaude 3 Opus by 7.290.783.5+7.2
- MMLU 5-shot CoTClaude 3 Opus by 6.788.281.5+6.7
- XSTestClaude 3 Opus by 6.792.585.8+6.7
- HellaSwagClaude 3 Opus by 6.495.489+6.4
- MMMU (validation)Claude 3 Opus by 6.359.453.1+6.3
- BBQ Ambig Accuracy (%)Claude 3 Opus by 598.693.6+5
- DROPClaude 3 Opus by 4.283.178.9+4.2
- DROP F1 ScoreClaude 3 Opus by 4.283.178.9+4.2
- bbqClaude 3 Opus by 49490+4
- BBHClaude 3 Opus by 3.986.882.9+3.9
- BIG-Bench Hard 3-shot CoTClaude 3 Opus by 3.986.882.9+3.9
- BBQ Ambig Bias (%)Claude 3 Sonnet by 3.81.25+3.8
- OpenBookQAClaude 3 Opus by 3.895.691.8+3.8
- ARC-ChallengeClaude 3 Opus by 3.296.493.2+3.2
- AA IntelligenceClaude 3 Opus by 2.88.75.9+2.8
- GSM8KClaude 3 Opus by 2.79592.3+2.7
- HarmBenchClaude 3 Opus by 1.697.495.8+1.6
- BBQ Disambig Bias (%)tie0.81.2tie
- AIR-Bench 2024tie84.484.7tie
- DocVQA (test, ANLS score)tie89.389.5tie
- anthropic_red_teamtie99.899.8tie
- simple_safety_teststie100100tie
Questions people ask
Which is better, Claude 3 Opus or Claude 3 Sonnet?
Claude 3 Opus wins two of the three areas where both have results: reasoning and images and charts. Claude 3 Sonnet wins none. They are level on coding.
Which is better for coding?
Neither. All 1 coding tests both models report are ties.
How do you compare the two?
We use the 38 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 34 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.