Claude 3.5 Sonnet vs GPT-4o
Wins 2 of 4 areas
Coding · Reasoning
Wins 2 of 4 areas
Images and charts · Long documents
The two are evenly matched.Claude 3.5 Sonnet is better at reasoning and coding; GPT-4o at images and charts and long documents.
Scores updated · 106 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Claude 3.5 Sonnet costs 10% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where GPT-4o pulls ahead
- Questions about very long textsLongBench v2+7.1points ahead
- College exam questions with charts, maps and diagramsMMMU+3.9points ahead
- Code for real scientific research problemsSciCode+1.7points ahead
Where Claude 3.5 Sonnet pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+15.8points ahead
- Recent programming contest problemsLiveCodeBench v6+6.3points ahead
- Math problems shown in pictures and chartsMathVista+6.3points ahead
Every test, side by side
All 106 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude 3.5 Sonnet
- SWE-bench VerifiedClaude 3.5 Sonnet by 15.84933.2+15.8
- LiveCodeBench v6Claude 3.5 Sonnet by 6.337.230.9+6.3
- SciCodeGPT-4o by 1.731.633.3+1.7
ReasoningClaude 3.5 Sonnet
- GPQA DiamondClaude 3.5 Sonnet by 3.45652.6+3.4
- Humanity's Last ExamClaude 3.5 Sonnet by 1.43.21.8+1.4
Images and chartsGPT-4o
Other results95 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MME_sumGPT-4o by 409 rating points19202329+409 rating
- VCR_en easyGPT-4o by 27.763.991.5+27.7
- AIR-Bench 2024Claude 3.5 Sonnet by 23.585.962.4+23.5
- metr_hcastClaude 3.5 Sonnet by 15.740.725+15.7
- RealWorldQAGPT-4o by 15.360.175.4+15.3
- HarmBenchClaude 3.5 Sonnet by 15.298.182.9+15.2
- MMLUClaude 3.5 Sonnet by 14.988.773.8+14.9
- Aider-PolyglotClaude 3.5 Sonnet by 14.645.330.7+14.6
- MATHGPT-4o by 14.271.185.3+14.2
- LongBench v2 medium (w/o CoT)GPT-4o by 13.838.652.4+13.8
- Aider-Edit (Acc.)Claude 3.5 Sonnet by 11.384.272.9+11.3
- LongBench v2 easy (w/o CoT)GPT-4o by 10.546.957.4+10.5
- MTOB eng → kalam (ChrF) no contextClaude 3.5 Sonnet by 10.320.29.9+10.3
- M-LongDoc (multimodal long-document benchmark)GPT-4o by 1031.441.4+10
- SimpleQAGPT-4o by 9.828.438.2+9.8
- LongBench v2 overall (w/o CoT)GPT-4o by 9.14150.1+9.1
- TAU-bench (retail)Claude 3.5 Sonnet by 8.969.260.3+8.9
- AIME 2025GPT-4o by 8.43.311.7+8.4
- LongBench v2 hard (w/o CoT)GPT-4o by 8.337.345.6+8.3
- LongBench v2 hard (w/ CoT) (Pass@1 with chain-of-thought)GPT-4o by 8.241.549.7+8.2
- FRAMES (Acc.)GPT-4o by 872.580.5+8
- LongBench v2 short (w/o CoT)GPT-4o by 7.246.153.3+7.2
- RULER 64KClaude 3.5 Sonnet by 6.895.288.4+6.8
- AIME 2024Claude 3.5 Sonnet by 6.7169.3+6.7
- LongBench v2 medium (w/ CoT) (Pass@1 with chain-of-thought)GPT-4o by 6.741.948.6+6.7
- Ruler 16kClaude 3.5 Sonnet by 6.795.789+6.7
- Ruler 32kClaude 3.5 Sonnet by 6.29588.8+6.2
- IFEval (avg)Claude 3.5 Sonnet by 690.184.1+6
- GPQA (Diamond) 0-shot CoTClaude 3.5 Sonnet by 5.859.453.6+5.8
- SuperGPQAClaude 3.5 Sonnet by 5.848.242.4+5.8
- LongBench v2 short (w/ CoT) (Pass@1 with chain-of-thought)GPT-4o by 5.753.959.6+5.7
- GSM8KClaude 3.5 Sonnet by 5.596.490.9+5.5
- IFEvalClaude 3.5 Sonnet by 5.586.581+5.5
- LiveCodeBenchGPT-4o by 5.532.838.3+5.5
- AlpacaEval 2 LCGPT-4o by 5.152.457.5+5.1
- ChartQAClaude 3.5 Sonnet by 5.190.885.7+5.1
- HallBench_avgGPT-4o by 5.149.955+5.1
- metr_re_benchClaude 3.5 Sonnet by 550+5
- NarrativeQAGPT-4o by 4.974.679.5+4.9
- Arena-Hard (GPT-4-1106 judge)Claude 3.5 Sonnet by 4.885.280.4+4.8
- LongBench v2 overall (w/ CoT) (Pass@1 with chain-of-thought)GPT-4o by 4.746.751.4+4.7
- DROP (3-shot F1)Claude 3.5 Sonnet by 4.688.383.7+4.6
- CodeforcesGPT-4o by 42 rating points717759+42 rating
- Codeforces (Rating)GPT-4o by 42 rating points717759+42 rating
- Ruler 8kClaude 3.5 Sonnet by 3.99692.1+3.9
- DROPClaude 3.5 Sonnet by 3.787.183.4+3.7
- MATH-500 (EM)Claude 3.5 Sonnet by 3.778.374.6+3.7
- MATH500 (Pass@1)Claude 3.5 Sonnet by 3.778.374.6+3.7
- MMBench-EN_testGPT-4o by 3.779.783.4+3.7
- MMBench-V1.1_testGPT-4o by 3.778.582.2+3.7
- HumanEvalClaude 3.5 Sonnet by 3.593.790.2+3.5
- Chinese SimpleQA (C-SimpleQA)GPT-4o by 3.355.458.7+3.3
- Codeforces (Percentile)GPT-4o by 3.320.323.6+3.3
- LongBench v2 long (w/o CoT)GPT-4o by 3.23740.2+3.2
- OlympiadBenchClaude 3.5 Sonnet by 3.228.425.2+3.2
- TAU-bench (airline)Claude 3.5 Sonnet by 3.24642.8+3.2
- MMLU-ProClaude 3.5 Sonnet by 2.977.674.7+2.9
- ChartQA_relaxedClaude 3.5 Sonnet by 2.790.888.1+2.7
- CLUEWSCGPT-4o by 2.585.487.9+2.5
- DocVQAClaude 3.5 Sonnet by 2.495.292.8+2.4
- DocVQA (test, ANLS score)Claude 3.5 Sonnet by 2.495.292.8+2.4
- DocVQA_testClaude 3.5 Sonnet by 2.495.292.8+2.4
- CNMO 2024Claude 3.5 Sonnet by 2.313.110.8+2.3
- MEGA-Bench_macroClaude 3.5 Sonnet by 251.449.4+2
- Artificial Analysis Coding IndexClaude 3.5 Sonnet by 1.82624.2+1.8
- HumanEval 0-shotClaude 3.5 Sonnet by 1.89290.2+1.8
- MTOB kalam → eng (BLEURT) no contextGPT-4o by 1.831.433.2+1.8
- XSTestGPT-4o by 1.795.697.3+1.7
- simple_safety_testsClaude 3.5 Sonnet by 1.510098.5+1.5
- MMBench-CN_testGPT-4o by 1.480.782.1+1.4
- MTOB kalam → eng (BLEURT) half bookClaude 3.5 Sonnet by 1.459.758.3+1.4
- HumanEval-Mul (Pass@1)Claude 3.5 Sonnet by 1.281.780.5+1.2
- MBPP+ (EvalPlus-augmented)GPT-4o by 1.175.176.2+1.1
- MGSMClaude 3.5 Sonnet by 1.191.690.5+1.1
- MGSM 0-shot CoTClaude 3.5 Sonnet by 1.191.690.5+1.1
- LongBench v2 easy (w/ CoT) (Pass@1 with chain-of-thought)Claude 3.5 Sonnet by 155.254.2+1
- AlpacaEval2.0 (LC-winrate)tie5251.1tie
- MMLU-Reduxtie88.988tie
- MMMU (val) (Pass@1)tie68.369.1tie
- MMMU (validation)tie68.369.1tie
- anthropic_red_teamtie99.899.1tie
- C-Evaltie76.776tie
- MTOB eng → kalam (ChrF) half booktie53.654.3tie
- naturalquestions_closedbooktie50.249.6tie
- AI2Dtie94.794.2tie
- metr_swaatie98.799.2tie
- Ruler 4ktie96.597tie
- DROP (F1)tie88.889.2tie
- LiveCodeBench (Pass@1-COT)tie33.834.2tie
- MMLU 0-shot CoTtie88.388.7tie
- OpenBookQAtie97.296.8tie
- bbqtie94.995.1tie
- AA Intelligencetie7.27.3tie
- Arena Hardtie79.279.3tie
- MT-Benchtie8.88.7tie
Questions people ask
Which is better, Claude 3.5 Sonnet or GPT-4o?
Claude 3.5 Sonnet and GPT-4o each win two of the four areas where both have results. Claude 3.5 Sonnet wins coding and reasoning; GPT-4o wins images and charts and long documents. Claude 3.5 Sonnet costs 10% less.
Which is better for coding?
Claude 3.5 Sonnet. It wins 2 of the 3 coding tests both models report; GPT-4o wins 1.
Which is cheaper?
Claude 3.5 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens; GPT-4o costs $5.00 and $15.00. That makes Claude 3.5 Sonnet about 10% cheaper for the same work.
How do you compare the two?
We use the 106 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 95 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.