DeepSeek-R1-Distill-Qwen-14B vs GPT-4o
Wins 0 of 4 areas
—
Wins 2 of 4 areas
Long documents · Following instructions
GPT-4o wins more areas, narrowly.DeepSeek-R1-Distill-Qwen-14B is cheaper.
Scores updated · 17 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-R1-Distill-Qwen-14B costs 98% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4o pulls ahead
Where DeepSeek-R1-Distill-Qwen-14B pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+2.3points ahead
- Recent programming contest problemsLiveCodeBench v6+1points ahead
Every test, side by side
All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- SciCodeGPT-4o by 9.423.933.3+9.4
- LiveCodeBench v6DeepSeek-R1-Distill-Qwen-14B by 131.930.9+1
ReasoningEven
- GPQA DiamondGPT-4o by 4.248.452.6+4.2
- Humanity's Last ExamDeepSeek-R1-Distill-Qwen-14B by 2.34.11.8+2.3
Other results11 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforces (Rating)DeepSeek-R1-Distill-Qwen-14B by 722 rating points1481759+722 rating
- AIME 2024DeepSeek-R1-Distill-Qwen-14B by 60.469.79.3+60.4
- AIME 2025DeepSeek-R1-Distill-Qwen-14B by 37.549.211.7+37.5
- MATH-500 (EM)DeepSeek-R1-Distill-Qwen-14B by 19.393.974.6+19.3
- MATH500 (Pass@1)DeepSeek-R1-Distill-Qwen-14B by 19.393.974.6+19.3
- LiveCodeBenchDeepSeek-R1-Distill-Qwen-14B by 14.853.138.3+14.8
- MMLU-ProGPT-4o by 5.968.874.7+5.9
- IFEvalGPT-4o by 2.778.381+2.7
- DROP (3-shot F1)DeepSeek-R1-Distill-Qwen-14B by 1.885.583.7+1.8
- SuperGPQAGPT-4o by 1.840.642.4+1.8
- AA Intelligencetie7.87.3tie
Questions people ask
Which is better, DeepSeek-R1-Distill-Qwen-14B or GPT-4o?
GPT-4o wins two of the four areas where both have results: long documents and following instructions. DeepSeek-R1-Distill-Qwen-14B wins none, but costs 98% less. They are level on coding and reasoning.
Which is better for coding?
Neither. They win 1 coding test each of the 2 both models report.
Which is cheaper?
DeepSeek-R1-Distill-Qwen-14B costs $0.20 per million input tokens and $0.20 per million output tokens; GPT-4o costs $5.00 and $15.00. That makes DeepSeek-R1-Distill-Qwen-14B about 98% cheaper for the same work.
How do you compare the two?
We use the 17 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 11 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.