DeepSeek-R1 0528 vs DeepSeek-V3.1
Wins 2 of 6 areas
Reasoning · Facts
Wins 3 of 6 areas
Coding · Long documents · Following instructions
DeepSeek-V3.1 wins more areas, narrowly.DeepSeek-R1 0528 is better at reasoning.
Scores updated · 32 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software13DeepSeek-V3.13 of 4 tests
- Following instructionsDoing exactly what it is asked02DeepSeek-V3.12 of 2 tests
- Long documentsFinding answers in very long texts01DeepSeek-V3.11 of 1 test
- ReasoningHard problems that need careful thinking20DeepSeek-R1 05282 of 4 tests · 2 ties
- FactsGetting facts right instead of making them up10DeepSeek-R1 05281 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own11Even1 each
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V3.1 costs 49% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where DeepSeek-V3.1 pulls ahead
- Fixes real GitHub issues in many programming languagesSWE-bench Multilingual+24points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+21.4points ahead
- Finds hard-to-locate facts by browsing the webBrowseComp+21.1points ahead
Where DeepSeek-R1 0528 pulls ahead
- Real work tasks from 44 professionsGDPVal+3.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+3.4points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.5points ahead
Every test, side by side
All 32 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V3.1
- SWE-bench MultilingualDeepSeek-V3.1 by 2430.554.5+24
- SWE-bench VerifiedDeepSeek-V3.1 by 21.444.666+21.4
- Terminal-Bench HardDeepSeek-V3.1 by 9.115.925+9.1
- SciCodeDeepSeek-R1 0528 by 1.240.339.1+1.2
AgentsEven
- BrowseCompDeepSeek-V3.1 by 21.18.930+21.1
- GDPValDeepSeek-R1 0528 by 3.495.6+3.4
ReasoningDeepSeek-R1 0528
- GPQA DiamondDeepSeek-R1 0528 by 3.481.377.9+3.4
- Humanity's Last ExamDeepSeek-R1 0528 by 1.515.814.3+1.5
- SimpleBenchtie40.840tie
- CritPttie1.42tie
FactsDeepSeek-R1 0528
- AA-Omniscience · AccuracyDeepSeek-R1 0528 by 1.530.529+1.5
- AA-Omniscience · Non-hallucinationtie16.617.5tie
Following instructionsDeepSeek-V3.1
- IFBenchDeepSeek-V3.1 by 1.939.641.5+1.9
- Multi-ChallengeDeepSeek-V3.1 by 1.14546.1+1.1
Other results17 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HMMT 2025DeepSeek-R1 0528 by 45.979.433.5+45.9
- AIME 2025DeepSeek-R1 0528 by 37.787.549.8+37.7
- Terminal-BenchDeepSeek-V3.1 by 25.65.731.3+25.6
- LiveCodeBenchDeepSeek-R1 0528 by 16.973.356.4+16.9
- Codeforces-Div1 (Rating)DeepSeek-V3.1 by 161 rating points19302091+161 rating
- BrowseComp-ZHDeepSeek-V3.1 by 13.535.749.2+13.5
- HMMT Feb. 2025DeepSeek-V3.1 by 9.176.785.8+9.1
- Artificial Analysis Coding IndexDeepSeek-V3.1 by 5.72429.7+5.7
- Aider-PolyglotDeepSeek-R1 0528 by 3.271.668.4+3.2
- AA-OmniscienceDeepSeek-R1 0528 by 2.2-27.4-29.6+2.2
- AA Agentic IndexDeepSeek-R1 0528 by 1.920.818.9+1.9
- AIME 2024DeepSeek-V3.1 by 1.791.493.1+1.7
- MMLU-ReduxDeepSeek-R1 0528 by 1.693.491.8+1.6
- MMLU-ProDeepSeek-R1 0528 by 1.38583.7+1.3
- SimpleQADeepSeek-V3.1 by 1.192.393.4+1.1
- τ²-Bench Telecom (AA run)tie36.537.4tie
- AA Intelligencetie13.113.5tie
Questions people ask
Which is better, DeepSeek-R1 0528 or DeepSeek-V3.1?
DeepSeek-V3.1 wins three of the six areas where both have results: coding, long documents and following instructions. DeepSeek-R1 0528 wins reasoning and facts. They are level on agents.
Which is better for coding?
DeepSeek-V3.1. It wins 3 of the 4 coding tests both models report; DeepSeek-R1 0528 wins 1.
Which is cheaper?
DeepSeek-R1 0528 costs $1.35 per million input tokens and $3.00 per million output tokens; DeepSeek-V3.1 costs $0.56 and $1.68. That makes DeepSeek-V3.1 about 49% cheaper for the same work.
How do you compare the two?
We use the 32 benchmark tests both models have published scores on. The verdict counts the 15 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 17 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.