The two are evenly matched.DeepSeek-R1 is better at coding; DeepSeek-R1-Zero at reasoning.
Scores updated · 18 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and reasoning rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-R1 pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+6points ahead
Where DeepSeek-R1-Zero pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+2.5points ahead
Every test, side by side
All 18 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results16 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AlpacaEval2.0 (LC-winrate)DeepSeek-R1 by 62.987.624.7+62.9
- CodeforcesDeepSeek-R1 by 585 rating points20291444+585 rating
- Aider-PolyglotDeepSeek-R1 by 41.153.312.2+41.1
- Arena-Hard (GPT-4-1106 judge)DeepSeek-R1 by 38.792.353.6+38.7
- IFEvalDeepSeek-R1 by 36.783.346.6+36.7
- MMLU-ProDeepSeek-R1 by 15.18468.9+15.1
- LiveCodeBenchDeepSeek-R1 by 13.563.550+13.5
- MMLU-ReduxDeepSeek-R1 by 7.392.985.6+7.3
- DROP (3-shot F1)DeepSeek-R1 by 3.192.289.1+3.1
- Chinese SimpleQA (C-SimpleQA)DeepSeek-R1-Zero by 2.763.766.4+2.7
- MMLUDeepSeek-R1 by 290.888.8+2
- AIME 2024DeepSeek-R1 by 1.979.877.9+1.9
- MATH-500 (EM)DeepSeek-R1 by 1.497.395.9+1.4
- CLUEWSCtie92.893.1tie
- FRAMES (Acc.)tie82.582.3tie
- SimpleQAtie30.130.3tie
Questions people ask
Which is better, DeepSeek-R1 or DeepSeek-R1-Zero?
DeepSeek-R1 and DeepSeek-R1-Zero each win one of the two areas where both have results. DeepSeek-R1 wins coding; DeepSeek-R1-Zero wins reasoning.
Which is better for coding?
DeepSeek-R1. It wins the one coding test both models report.
How do you compare the two?
We use the 18 benchmark tests both models have published scores on. The verdict counts the 2 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 16 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.