Claude Sonnet 4 vs DeepSeek-R1-Distill-Qwen-14B
Wins 4 of 4 areas
Coding · Reasoning · Long documents · Following instructions
Wins 0 of 4 areas
—
Claude Sonnet 4 is the stronger all-rounder.DeepSeek-R1-Distill-Qwen-14B is cheaper.
Scores updated · 14 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-R1-Distill-Qwen-14B costs 98% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude Sonnet 4 pulls ahead
- Reasons across sets of long documentsAA-LCR+60points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+32.6points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+29.3points ahead
Where DeepSeek-R1-Distill-Qwen-14B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 14 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Sonnet 4
- LiveCodeBench v6Claude Sonnet 4 by 16.648.531.9+16.6
- SciCodeClaude Sonnet 4 by 16.14023.9+16.1
ReasoningClaude Sonnet 4
- GPQA DiamondClaude Sonnet 4 by 29.377.748.4+29.3
- Humanity's Last ExamClaude Sonnet 4 by 6.610.74.1+6.6
Following instructionsClaude Sonnet 4
- IFBenchClaude Sonnet 4 by 32.654.722.1+32.6
Other results8 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AIME 2024DeepSeek-R1-Distill-Qwen-14B by 26.343.469.7+26.3
- AIME 2025Claude Sonnet 4 by 21.370.549.2+21.3
- LiveCodeBenchClaude Sonnet 4 by 15.468.553.1+15.4
- SuperGPQAClaude Sonnet 4 by 15.155.740.6+15.1
- MMLU-ProClaude Sonnet 4 by 14.983.768.8+14.9
- AA IntelligenceClaude Sonnet 4 by 11.118.97.8+11.1
- IFEvalClaude Sonnet 4 by 9.387.678.3+9.3
- MATH-500 (EM)tie9493.9tie
Questions people ask
Which is better, Claude Sonnet 4 or DeepSeek-R1-Distill-Qwen-14B?
Claude Sonnet 4 wins all four areas where both have results: coding, reasoning, long documents and following instructions. DeepSeek-R1-Distill-Qwen-14B wins none, but costs 98% less.
Which is better for coding?
Claude Sonnet 4. It wins 2 of the 2 coding tests both models report; DeepSeek-R1-Distill-Qwen-14B wins none.
Which is cheaper?
Claude Sonnet 4 costs $3.00 per million input tokens and $15.00 per million output tokens; DeepSeek-R1-Distill-Qwen-14B costs $0.20 and $0.20. That makes DeepSeek-R1-Distill-Qwen-14B about 98% cheaper for the same work.
How do you compare the two?
We use the 14 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 8 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.