Claude 3.5 Sonnet vs DeepSeek-V3.2
Wins 0 of 3 areas
—
Wins 3 of 3 areas
Coding · Reasoning · Long documents
DeepSeek-V3.2 is the stronger all-rounder.
Scores updated · 14 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V3.2 costs 96% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3.2 pulls ahead
- Recent programming contest problemsLiveCodeBench v6+46.1points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+28points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+24.1points ahead
Where Claude 3.5 Sonnet pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 14 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V3.2
- LiveCodeBench v6DeepSeek-V3.2 by 46.137.283.3+46.1
- SWE-bench VerifiedDeepSeek-V3.2 by 24.14973.1+24.1
- SciCodeDeepSeek-V3.2 by 7.331.638.9+7.3
ReasoningDeepSeek-V3.2
- GPQA DiamondDeepSeek-V3.2 by 285684+28
- Humanity's Last ExamDeepSeek-V3.2 by 21.43.224.6+21.4
Other results8 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HMMT Feb. 2025DeepSeek-V3.2 by 90.81.792.5+90.8
- AIME 2025DeepSeek-V3.2 by 89.83.393.1+89.8
- LiveCodeBenchDeepSeek-V3.2 by 50.532.883.3+50.5
- Artificial Analysis Coding IndexDeepSeek-V3.2 by 18.22644.2+18.2
- AA IntelligenceDeepSeek-V3.2 by 14.37.221.5+14.3
- MMLU-ProDeepSeek-V3.2 by 7.477.685+7.4
- MMLU-ReduxDeepSeek-V3.2 by 4.888.993.7+4.8
- frontiermath_tier_4_v1DeepSeek-V3.2 by 2.102.1+2.1
Questions people ask
Which is better, Claude 3.5 Sonnet or DeepSeek-V3.2?
DeepSeek-V3.2 wins all three areas where both have results: coding, reasoning and long documents. Claude 3.5 Sonnet wins none.
Which is better for coding?
DeepSeek-V3.2. It wins 3 of the 3 coding tests both models report; Claude 3.5 Sonnet wins none.
Which is cheaper?
Claude 3.5 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens; DeepSeek-V3.2 costs $0.28 and $0.42. That makes DeepSeek-V3.2 about 96% cheaper for the same work.
How do you compare the two?
We use the 14 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 8 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.