GPT-4o is the stronger all-rounder.
Scores updated · 25 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding, reasoning and long documents rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4o pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+16.4points ahead
- Questions about very long textsLongBench v2+12.7points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+10.3points ahead
Where DeepSeek-V2.5 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 25 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- CodeforcesGPT-4o by 723.4 rating points35.6759+723.4 rating
- SimpleQAGPT-4o by 2810.238.2+28
- FRAMES (Acc.)GPT-4o by 15.165.480.5+15.1
- Aider-PolyglotGPT-4o by 12.917.830.7+12.9
- MATHGPT-4o by 10.674.785.3+10.6
- MMLU-ProGPT-4o by 8.566.274.7+8.5
- MMLU-ReduxGPT-4o by 7.780.388+7.7
- AIME 2024DeepSeek-V2.5 by 7.416.79.3+7.4
- HumanEval-Mul (Pass@1)GPT-4o by 6.773.880.5+6.7
- MMLUDeepSeek-V2.5 by 6.680.473.8+6.6
- Chinese SimpleQA (C-SimpleQA)GPT-4o by 4.654.158.7+4.6
- GSM8KDeepSeek-V2.5 by 4.295.190.9+4.2
- DROP (3-shot F1)DeepSeek-V2.5 by 4.187.883.7+4.1
- LiveCodeBenchDeepSeek-V2.5 by 3.541.838.3+3.5
- Arena HardGPT-4o by 3.176.279.3+3.1
- CLUEWSCDeepSeek-V2.5 by 2.590.487.9+2.5
- Aider-Edit (Acc.)GPT-4o by 1.371.672.9+1.3
- HumanEvalGPT-4o by 1.28990.2+1.2
- AA Intelligencetie6.67.3tie
- IFEvaltie80.681tie
- MT-Benchtie98.7tie
- MATH-500 (EM)tie74.774.6tie
Questions people ask
Which is better, DeepSeek-V2.5 or GPT-4o?
GPT-4o wins all three areas where both have results: coding, reasoning and long documents. DeepSeek-V2.5 wins none.
Which is better for coding?
GPT-4o. It wins the one coding test both models report.
How do you compare the two?
We use the 25 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.