DeepSeek-R1-Distill-Qwen-14B vs DeepSeek-V3
Wins 0 of 4 areas
—
Wins 4 of 4 areas
Coding · Reasoning · Long documents · Following instructions
DeepSeek-V3 is the stronger all-rounder.DeepSeek-R1-Distill-Qwen-14B is cheaper.
Scores updated · 17 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-R1-Distill-Qwen-14B costs 65% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V3 pulls ahead
- Reasons across sets of long documentsAA-LCR+30.4points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+18.9points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+17.1points ahead
Where DeepSeek-R1-Distill-Qwen-14B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V3
- SciCodeDeepSeek-V3 by 15.123.939+15.1
- LiveCodeBench v6DeepSeek-V3 by 1531.946.9+15
ReasoningDeepSeek-V3
- GPQA DiamondDeepSeek-V3 by 17.148.465.5+17.1
- Humanity's Last Examtie4.14.7tie
Following instructionsDeepSeek-V3
- IFBenchDeepSeek-V3 by 18.922.141+18.9
Other results11 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforces (Rating)DeepSeek-R1-Distill-Qwen-14B by 347 rating points14811134+347 rating
- AIME 2024DeepSeek-R1-Distill-Qwen-14B by 30.569.739.2+30.5
- SuperGPQADeepSeek-V3 by 13.140.653.7+13.1
- IFEvalDeepSeek-V3 by 7.878.386.1+7.8
- MMLU-ProDeepSeek-V3 by 7.168.875.9+7.1
- DROP (3-shot F1)DeepSeek-V3 by 6.185.591.6+6.1
- LiveCodeBenchDeepSeek-R1-Distill-Qwen-14B by 3.953.149.2+3.9
- MATH-500 (EM)DeepSeek-R1-Distill-Qwen-14B by 3.793.990.2+3.7
- HMMT Feb. 2025DeepSeek-R1-Distill-Qwen-14B by 2.531.729.2+2.5
- AIME 2025DeepSeek-V3 by 2.149.251.3+2.1
- AA IntelligenceDeepSeek-V3 by 1.97.89.7+1.9
Questions people ask
Which is better, DeepSeek-R1-Distill-Qwen-14B or DeepSeek-V3?
DeepSeek-V3 wins all four areas where both have results: coding, reasoning, long documents and following instructions. DeepSeek-R1-Distill-Qwen-14B wins none, but costs 65% less.
Which is better for coding?
DeepSeek-V3. It wins 2 of the 2 coding tests both models report; DeepSeek-R1-Distill-Qwen-14B wins none.
Which is cheaper?
DeepSeek-R1-Distill-Qwen-14B costs $0.20 per million input tokens and $0.20 per million output tokens; DeepSeek-V3 costs $0.24 and $0.90. That makes DeepSeek-R1-Distill-Qwen-14B about 65% cheaper for the same work.
How do you compare the two?
We use the 17 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 11 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.