Command R+ (Apr '24) wins more areas, narrowly.
Scores updated · 16 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Command R+ (Apr '24) pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.7points ahead
Where DBRX Instruct pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 16 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningCommand R+ (Apr '24)
- Humanity's Last ExamCommand R+ (Apr '24) by 1.74.62.9+1.7
- GPQA Diamondtie32.333.1tie
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- simple_safety_testsCommand R+ (Apr '24) by 46.510053.5+46.5
- NarrativeQACommand R+ (Apr '24) by 24.773.548.8+24.7
- anthropic_red_teamCommand R+ (Apr '24) by 21.49876.6+21.4
- HarmBenchCommand R+ (Apr '24) by 21.448.527.1+21.4
- XSTestCommand R+ (Apr '24) by 16.493.877.4+16.4
- bbqCommand R+ (Apr '24) by 10.789.979.2+10.7
- OpenBookQADBRX Instruct by 8.282.891+8.2
- naturalquestions_closedbookCommand R+ (Apr '24) by 5.934.328.4+5.9
- MMLUDBRX Instruct by 5.35964.3+5.3
- MATHCommand R+ (Apr '24) by 4.540.335.8+4.5
- AIR-Bench 2024Command R+ (Apr '24) by 3.929.325.4+3.9
- GSM8KCommand R+ (Apr '24) by 3.670.767.1+3.6
- AA Intelligencetie5.35.3tie
Questions people ask
Which is better, Command R+ (Apr '24) or DBRX Instruct?
Command R+ (Apr '24) wins one of the two areas where both have results: reasoning. DBRX Instruct wins none. They are level on coding.
Which is better for coding?
Neither. All 1 coding tests both models report are ties.
How do you compare the two?
We use the 16 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.