GPT-4o is the stronger all-rounder.DeepSeek-R1-Distill-Qwen-7B is cheaper.
Scores updated · 17 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and reasoning rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-R1-Distill-Qwen-7B costs 99% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4o pulls ahead
- Recent programming contest problemsLiveCodeBench v6+7points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+3.5points ahead
Where DeepSeek-R1-Distill-Qwen-7B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 17 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results15 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Arena HardGPT-4o by 68.910.479.3+68.9
- AIME 2024DeepSeek-R1-Distill-Qwen-7B by 46.255.59.3+46.2
- Codeforces (Rating)DeepSeek-R1-Distill-Qwen-7B by 430 rating points1189759+430 rating
- DROPGPT-4o by 31.651.883.4+31.6
- AIME 2025DeepSeek-R1-Distill-Qwen-7B by 27.138.811.7+27.1
- MMLUGPT-4o by 23.150.773.8+23.1
- MMLU-ProGPT-4o by 21.253.574.7+21.2
- IFEvalGPT-4o by 20.560.581+20.5
- MATH-500 (EM)DeepSeek-R1-Distill-Qwen-7B by 18.292.874.6+18.2
- MATH500 (Pass@1)DeepSeek-R1-Distill-Qwen-7B by 18.292.874.6+18.2
- SuperGPQAGPT-4o by 13.528.942.4+13.5
- GSM8KGPT-4o by 12.478.590.9+12.4
- HumanEvalGPT-4o by 10.379.990.2+10.3
- DROP (3-shot F1)GPT-4o by 6.77783.7+6.7
- LiveCodeBenchtie37.638.3tie
Questions people ask
Which is better, DeepSeek-R1-Distill-Qwen-7B or GPT-4o?
GPT-4o wins all two areas where both have results: coding and reasoning. DeepSeek-R1-Distill-Qwen-7B wins none, but costs 99% less.
Which is better for coding?
GPT-4o. It wins the one coding test both models report.
Which is cheaper?
DeepSeek-R1-Distill-Qwen-7B costs $0.15 per million input tokens and $0.15 per million output tokens; GPT-4o costs $5.00 and $15.00. That makes DeepSeek-R1-Distill-Qwen-7B about 99% cheaper for the same work.
How do you compare the two?
We use the 17 benchmark tests both models have published scores on. The verdict counts the 2 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 15 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.