GPT-4o is the stronger all-rounder.
Scores updated · 28 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding and images and charts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4o pulls ahead
- Math problems shown in pictures and chartsMathVista+13.5points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+12.6points ahead
- Code for real scientific research problemsSciCode+10.4points ahead
Where Claude 3 Sonnet pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.8points ahead
Every test, side by side
All 28 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningEven
- GPQA DiamondGPT-4o by 12.64052.6+12.6
- Humanity's Last ExamClaude 3 Sonnet by 1.83.61.8+1.8
Other results24 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- NarrativeQAGPT-4o by 68.411.179.5+68.4
- naturalquestions_closedbookGPT-4o by 46.82.849.6+46.8
- MATHGPT-4o by 42.243.185.3+42.2
- AIR-Bench 2024Claude 3 Sonnet by 22.384.762.4+22.3
- MMLU-ProGPT-4o by 17.956.874.7+17.9
- HumanEvalGPT-4o by 17.27390.2+17.2
- HumanEval 0-shotGPT-4o by 17.27390.2+17.2
- MMMU (validation)GPT-4o by 1653.169.1+16
- GPQA (Diamond) 0-shot CoTGPT-4o by 13.240.453.6+13.2
- HarmBenchClaude 3 Sonnet by 12.995.882.9+12.9
- MMLU 0-shot CoTGPT-4o by 11.677.188.7+11.6
- XSTestGPT-4o by 11.585.897.3+11.5
- MGSMGPT-4o by 783.590.5+7
- MGSM 0-shot CoTGPT-4o by 783.590.5+7
- bbqGPT-4o by 5.19095.1+5.1
- OpenBookQAGPT-4o by 591.896.8+5
- DROPGPT-4o by 4.578.983.4+4.5
- DROP F1 ScoreGPT-4o by 4.578.983.4+4.5
- MMLUClaude 3 Sonnet by 4.578.373.8+4.5
- DocVQA (test, ANLS score)GPT-4o by 3.389.592.8+3.3
- simple_safety_testsClaude 3 Sonnet by 1.510098.5+1.5
- AA IntelligenceGPT-4o by 1.45.97.3+1.4
- GSM8KClaude 3 Sonnet by 1.492.390.9+1.4
- anthropic_red_teamtie99.899.1tie
Questions people ask
Which is better, Claude 3 Sonnet or GPT-4o?
GPT-4o wins two of the three areas where both have results: coding and images and charts. Claude 3 Sonnet wins none. They are level on reasoning.
Which is better for coding?
GPT-4o. It wins the one coding test both models report.
How do you compare the two?
We use the 28 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 24 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.