GPT-4o is the stronger all-rounder.
Scores updated · 35 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding, facts and images and charts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
GPT-4o costs 78% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-4o pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+13.4points ahead
- Math problems shown in pictures and chartsMathVista+10.9points ahead
- Code for real scientific research problemsSciCode+10points ahead
Where Claude 3 Opus pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1points ahead
Every test, side by side
All 35 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningEven
- GPQA DiamondGPT-4o by 3.748.952.6+3.7
- Humanity's Last ExamClaude 3 Opus by 12.81.8+1
Other results30 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- NarrativeQAGPT-4o by 44.435.179.5+44.4
- MATHGPT-4o by 25.260.185.3+25.2
- AIR-Bench 2024Claude 3 Opus by 2284.462.4+22
- HarmBenchClaude 3 Opus by 14.597.482.9+14.5
- MMLUClaude 3 Opus by 1386.873.8+13
- MMMU (validation)GPT-4o by 9.759.469.1+9.7
- MMLU-ProGPT-4o by 6.268.574.7+6.2
- AI2DGPT-4o by 6.188.194.2+6.1
- naturalquestions_closedbookGPT-4o by 5.54449.6+5.5
- HumanEvalGPT-4o by 5.384.990.2+5.3
- HumanEval 0-shotGPT-4o by 5.384.990.2+5.3
- XSTestGPT-4o by 4.892.597.3+4.8
- Artificial Analysis Coding IndexGPT-4o by 4.719.524.2+4.7
- metr_hcastGPT-4o by 4.420.625+4.4
- GSM8KClaude 3 Opus by 4.19590.9+4.1
- DocVQA (test, ANLS score)GPT-4o by 3.589.392.8+3.5
- GPQA (Diamond) 0-shot CoTGPT-4o by 3.250.453.6+3.2
- MMLU 0-shot CoTGPT-4o by 385.788.7+3
- Avg.GPT-4o by 2.125.727.8+2.1
- simple_safety_testsClaude 3 Opus by 1.510098.5+1.5
- AA IntelligenceClaude 3 Opus by 1.48.77.3+1.4
- OpenBookQAGPT-4o by 1.295.696.8+1.2
- bbqGPT-4o by 1.19495.1+1.1
- anthropic_red_teamtie99.899.1tie
- metr_swaatie98.599.2tie
- DROPtie83.183.4tie
- DROP F1 Scoretie83.183.4tie
- MGSMtie90.790.5tie
- MGSM 0-shot CoTtie90.790.5tie
- metr_re_benchtie00tie
Questions people ask
Which is better, Claude 3 Opus or GPT-4o?
GPT-4o wins three of the four areas where both have results: coding, facts and images and charts. Claude 3 Opus wins none. They are level on reasoning.
Which is better for coding?
GPT-4o. It wins the one coding test both models report.
Which is cheaper?
Claude 3 Opus costs $15.00 per million input tokens and $75.00 per million output tokens; GPT-4o costs $5.00 and $15.00. That makes GPT-4o about 78% cheaper for the same work.
How do you compare the two?
We use the 35 benchmark tests both models have published scores on. The verdict counts the 5 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 30 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.