Claude 3 Opus is the stronger all-rounder.Command R+ (Apr '24) is cheaper.
Scores updated · 19 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Command R+ (Apr '24) costs 86% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude 3 Opus pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+16.6points ahead
- Code for real scientific research problemsSciCode+11.5points ahead
- Common-sense trick questionsSimpleBench+6.1points ahead
Where Command R+ (Apr '24) pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.8points ahead
Every test, side by side
All 19 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningClaude 3 Opus
- GPQA DiamondClaude 3 Opus by 16.648.932.3+16.6
- SimpleBenchClaude 3 Opus by 6.123.517.4+6.1
- Humanity's Last ExamCommand R+ (Apr '24) by 1.82.84.6+1.8
Other results15 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIR-Bench 2024Claude 3 Opus by 55.184.429.3+55.1
- HarmBenchClaude 3 Opus by 48.997.448.5+48.9
- NarrativeQACommand R+ (Apr '24) by 38.435.173.5+38.4
- MMLUClaude 3 Opus by 27.886.859+27.8
- ARC-ChallengeClaude 3 Opus by 25.496.471+25.4
- GSM8KClaude 3 Opus by 24.39570.7+24.3
- MATHClaude 3 Opus by 19.860.140.3+19.8
- OpenBookQAClaude 3 Opus by 12.895.682.8+12.8
- naturalquestions_closedbookClaude 3 Opus by 9.84434.3+9.8
- HellaSwagClaude 3 Opus by 6.895.488.6+6.8
- bbqClaude 3 Opus by 4.19489.9+4.1
- AA IntelligenceClaude 3 Opus by 3.48.75.3+3.4
- anthropic_red_teamClaude 3 Opus by 1.899.898+1.8
- XSTestCommand R+ (Apr '24) by 1.392.593.8+1.3
- simple_safety_teststie100100tie
Questions people ask
Which is better, Claude 3 Opus or Command R+ (Apr '24)?
Claude 3 Opus wins all two areas where both have results: coding and reasoning. Command R+ (Apr '24) wins none, but costs 86% less.
Which is better for coding?
Claude 3 Opus. It wins the one coding test both models report.
Which is cheaper?
Claude 3 Opus costs $15.00 per million input tokens and $75.00 per million output tokens; Command R+ (Apr '24) costs $2.50 and $10.00. That makes Command R+ (Apr '24) about 86% cheaper for the same work.
How do you compare the two?
We use the 19 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 15 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.