DeepSeek-R1 vs Llama 3.1 Instruct 405B
Wins 3 of 6 areas
Coding · Reasoning · Long documents
Wins 0 of 6 areas
—
DeepSeek-R1 wins more areas, narrowly.
Scores updated · 37 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20DeepSeek-R12 of 3 tests · 1 tie
- ReasoningHard problems that need careful thinking20DeepSeek-R12 of 3 tests · 1 tie
- Long documentsFinding answers in very long texts20DeepSeek-R12 of 2 tests
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
- Following instructionsDoing exactly what it is asked00Even0 each · 1 tie
Agents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-R1 pulls ahead
- Reasons across sets of long documentsAA-LCR+32.4points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+24.7points ahead
- Questions about very long textsLongBench v2+22.2points ahead
Where Llama 3.1 Instruct 405B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+37.9points ahead
Every test, side by side
All 37 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-R1
- SWE-bench VerifiedDeepSeek-R1 by 24.749.224.5+24.7
- SciCodeDeepSeek-R1 by 8.438.329.9+8.4
- Terminal-Bench Hardtie6.16.8tie
ReasoningDeepSeek-R1
- GPQA DiamondDeepSeek-R1 by 19.370.851.5+19.3
- Humanity's Last ExamDeepSeek-R1 by 4.58.54+4.5
- CritPttie0.60tie
FactsEven
- AA-Omniscience · Non-hallucinationLlama 3.1 Instruct 405B by 37.99.747.6+37.9
- AA-Omniscience · AccuracyDeepSeek-R1 by 7.330.523.2+7.3
Long documentsDeepSeek-R1
- AA-LCRDeepSeek-R1 by 32.457.725.3+32.4
- LongBench v2DeepSeek-R1 by 22.258.336.1+22.2
Other results25 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- CodeforcesDeepSeek-R1 by 2004 rating points202925.3+2004 rating
- AIME 2024DeepSeek-R1 by 56.579.823.3+56.5
- LiveCodeBenchDeepSeek-R1 by 35.863.527.7+35.8
- MATHDeepSeek-R1 by 23.597.373.8+23.5
- MATH-500 (EM)DeepSeek-R1 by 23.597.373.8+23.5
- C-EvalDeepSeek-R1 by 19.391.872.5+19.3
- AA-OmniscienceLlama 3.1 Instruct 405B by 15.1-32.2-17.1+15.1
- HarmBenchLlama 3.1 Instruct 405B by 14.847.962.7+14.8
- Chinese SimpleQA (C-SimpleQA)DeepSeek-R1 by 13.363.750.4+13.3
- SimpleQADeepSeek-R1 by 1330.117.1+13
- FRAMES (Acc.)DeepSeek-R1 by 12.582.570+12.5
- MMLU-ProDeepSeek-R1 by 10.78473.3+10.7
- Artificial Analysis Coding IndexDeepSeek-R1 by 10.124.614.5+10.1
- CLUEWSCDeepSeek-R1 by 9.892.883+9.8
- τ²-Bench Telecom (AA run)Llama 3.1 Instruct 405B by 7.611.419+7.6
- MMLU-ReduxDeepSeek-R1 by 6.792.986.2+6.7
- AIR-Bench 2024Llama 3.1 Instruct 405B by 5.752.958.6+5.7
- IFEvalLlama 3.1 Instruct 405B by 5.383.388.6+5.3
- AA IntelligenceDeepSeek-R1 by 4.111.47.3+4.1
- DROP (3-shot F1)DeepSeek-R1 by 3.592.288.7+3.5
- MMLUDeepSeek-R1 by 2.290.888.6+2.2
- bbqDeepSeek-R1 by 2.196.694.5+2.1
- XSTestLlama 3.1 Instruct 405B by 1.594.495.9+1.5
- simple_safety_teststie9898.8tie
- anthropic_red_teamtie97.296.5tie
Questions people ask
Which is better, DeepSeek-R1 or Llama 3.1 Instruct 405B?
DeepSeek-R1 wins three of the six areas where both have results: coding, reasoning and long documents. Llama 3.1 Instruct 405B wins none. They are level on agents, facts and following instructions.
Which is better for coding?
DeepSeek-R1. It wins 2 of the 3 coding tests both models report; Llama 3.1 Instruct 405B wins none, and 1 is a tie.
How do you compare the two?
We use the 37 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 25 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.