DeepSeek-V3 is the stronger all-rounder.Command R is cheaper.
Scores updated · 16 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Command R costs 34% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3 pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+37.1points ahead
- Code for real scientific research problemsSciCode+32.7points ahead
Where Command R pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 16 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningDeepSeek-V3
- GPQA DiamondDeepSeek-V3 by 37.128.465.5+37.1
- Humanity's Last Examtie4.84.7tie
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MATHDeepSeek-V3 by 63.626.690.2+63.6
- GSM8KDeepSeek-V3 by 38.955.194+38.9
- MMLUDeepSeek-V3 by 31.856.788.5+31.8
- bbqDeepSeek-V3 by 24.372.496.7+24.3
- OpenBookQADeepSeek-V3 by 17.278.295.4+17.2
- naturalquestions_closedbookDeepSeek-V3 by 11.535.246.7+11.5
- AIR-Bench 2024DeepSeek-V3 by 931.840.8+9
- NarrativeQADeepSeek-V3 by 5.474.279.6+5.4
- AA IntelligenceDeepSeek-V3 by 4.759.7+4.7
- anthropic_red_teamDeepSeek-V3 by 3.493.797.1+3.4
- XSTestDeepSeek-V3 by 3.293.997.1+3.2
- simple_safety_testsDeepSeek-V3 by 194.395.3+1
- HarmBenchtie50.449.7tie
Questions people ask
Which is better, Command R or DeepSeek-V3?
DeepSeek-V3 wins all two areas where both have results: coding and reasoning. Command R wins none, but costs 34% less.
Which is better for coding?
DeepSeek-V3. It wins the one coding test both models report.
Which is cheaper?
Command R costs $0.15 per million input tokens and $0.60 per million output tokens; DeepSeek-V3 costs $0.24 and $0.90. That makes Command R about 34% cheaper for the same work.
How do you compare the two?
We use the 16 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.