Claude 3.5 Sonnet vs DeepSeek-V3
Wins 0 of 3 areas
—
Wins 3 of 3 areas
Coding · Reasoning · Long documents
DeepSeek-V3 is the stronger all-rounder.
Scores updated · 54 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V3 costs 94% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3 pulls ahead
- Recent programming contest problemsLiveCodeBench v6+9.7points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+9.5points ahead
- Questions about very long textsLongBench v2+7.7points ahead
Where Claude 3.5 Sonnet pulls ahead
- Common-sense trick questionsSimpleBench+8.6points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+7points ahead
Every test, side by side
All 54 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V3
- LiveCodeBench v6DeepSeek-V3 by 9.737.246.9+9.7
- SciCodeDeepSeek-V3 by 7.431.639+7.4
- SWE-bench VerifiedClaude 3.5 Sonnet by 74942+7
ReasoningDeepSeek-V3
- GPQA DiamondDeepSeek-V3 by 9.55665.5+9.5
- SimpleBenchClaude 3.5 Sonnet by 8.627.518.9+8.6
- Humanity's Last ExamDeepSeek-V3 by 1.53.24.7+1.5
Other results47 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HarmBenchClaude 3.5 Sonnet by 48.498.149.7+48.4
- AIME 2025DeepSeek-V3 by 483.351.3+48
- AIR-Bench 2024Claude 3.5 Sonnet by 45.185.940.8+45.1
- CodeforcesDeepSeek-V3 by 417 rating points7171134+417 rating
- Codeforces (Rating)DeepSeek-V3 by 417 rating points7171134+417 rating
- Codeforces (Percentile)DeepSeek-V3 by 38.420.358.7+38.4
- CNMO 2024DeepSeek-V3 by 30.113.143.2+30.1
- HumanEvalClaude 3.5 Sonnet by 28.593.765.2+28.5
- HMMT Feb. 2025DeepSeek-V3 by 27.51.729.2+27.5
- AIME 2024DeepSeek-V3 by 23.21639.2+23.2
- MATHDeepSeek-V3 by 19.171.190.2+19.1
- AlpacaEval2.0 (LC-winrate)DeepSeek-V3 by 185270+18
- LiveCodeBenchDeepSeek-V3 by 16.432.849.2+16.4
- Chinese SimpleQA (C-SimpleQA)DeepSeek-V3 by 12.655.468+12.6
- Arena HardDeepSeek-V3 by 12.279.291.4+12.2
- MATH-500 (EM)DeepSeek-V3 by 11.978.390.2+11.9
- MGSMClaude 3.5 Sonnet by 11.891.679.8+11.8
- C-EvalDeepSeek-V3 by 9.876.786.5+9.8
- LongBench v2 overall (w/o CoT)DeepSeek-V3 by 7.74148.7+7.7
- BBHClaude 3.5 Sonnet by 5.693.187.5+5.6
- CLUEWSCDeepSeek-V3 by 5.585.490.9+5.5
- SuperGPQADeepSeek-V3 by 5.548.253.7+5.5
- NarrativeQADeepSeek-V3 by 574.679.6+5
- simple_safety_testsClaude 3.5 Sonnet by 4.710095.3+4.7
- Aider-Edit (Acc.)Claude 3.5 Sonnet by 4.584.279.7+4.5
- DROPDeepSeek-V3 by 4.587.191.6+4.5
- Aider-PolyglotDeepSeek-V3 by 4.345.349.6+4.3
- MBPP+ (EvalPlus-augmented)DeepSeek-V3 by 3.775.178.8+3.7
- naturalquestions_closedbookClaude 3.5 Sonnet by 3.550.246.7+3.5
- SimpleQAClaude 3.5 Sonnet by 3.528.424.9+3.5
- DROP (3-shot F1)DeepSeek-V3 by 3.388.391.6+3.3
- Artificial Analysis Coding IndexClaude 3.5 Sonnet by 32623+3
- IFEval (avg)Claude 3.5 Sonnet by 2.890.187.3+2.8
- anthropic_red_teamClaude 3.5 Sonnet by 2.799.897.1+2.7
- AA IntelligenceDeepSeek-V3 by 2.57.29.7+2.5
- GSM8KClaude 3.5 Sonnet by 2.496.494+2.4
- bbqDeepSeek-V3 by 1.894.996.7+1.8
- OpenBookQAClaude 3.5 Sonnet by 1.897.295.4+1.8
- MMLU-ProClaude 3.5 Sonnet by 1.777.675.9+1.7
- XSTestDeepSeek-V3 by 1.595.697.1+1.5
- HumanEval-Mul (Pass@1)tie81.782.6tie
- FRAMES (Acc.)tie72.573.3tie
- IFEvaltie86.586.1tie
- Arena-Hard (GPT-4-1106 judge)tie85.285.5tie
- DROP (F1)tie88.889tie
- MMLUtie88.788.5tie
- MMLU-Reduxtie88.989.1tie
Questions people ask
Which is better, Claude 3.5 Sonnet or DeepSeek-V3?
DeepSeek-V3 wins all three areas where both have results: coding, reasoning and long documents. Claude 3.5 Sonnet wins none.
Which is better for coding?
DeepSeek-V3. It wins 2 of the 3 coding tests both models report; Claude 3.5 Sonnet wins 1.
Which is cheaper?
Claude 3.5 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens; DeepSeek-V3 costs $0.24 and $0.90. That makes DeepSeek-V3 about 94% cheaper for the same work.
How do you compare the two?
We use the 54 benchmark tests both models have published scores on. The verdict counts the 7 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 47 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.