Claude 3.5 Sonnet vs GPT-5.1
Wins 0 of 3 areas
—
Wins 3 of 3 areas
Coding · Reasoning · Images and charts
GPT-5.1 is the stronger all-rounder.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
GPT-5.1 costs 38% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-5.1 pulls ahead
- Recent programming contest problemsLiveCodeBench v6+49.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+31.3points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+27.3points ahead
Where Claude 3.5 Sonnet pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGPT-5.1
- LiveCodeBench v6GPT-5.1 by 49.837.287+49.8
- SWE-bench VerifiedGPT-5.1 by 27.34976.3+27.3
- SciCodeGPT-5.1 by 11.731.643.3+11.7
ReasoningGPT-5.1
- GPQA DiamondGPT-5.1 by 31.35687.3+31.3
- SimpleBenchGPT-5.1 by 25.727.553.2+25.7
- Humanity's Last ExamGPT-5.1 by 25.33.228.5+25.3
Images and chartsGPT-5.1
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HMMT Feb. 2025GPT-5.1 by 94.61.796.3+94.6
- AIME 2025GPT-5.1 by 90.73.394+90.7
- Artificial Analysis Coding IndexGPT-5.1 by 23.42649.4+23.4
- AA IntelligenceGPT-5.1 by 17.57.224.7+17.5
- MMMU (validation)GPT-5.1 by 17.168.385.4+17.1
- frontiermath_tier_4_v1GPT-5.1 by 12.5012.5+12.5
- MMLU-ProGPT-5.1 by 9.477.687+9.4
- bbqClaude 3.5 Sonnet by 6.294.988.7+6.2
- XSTestGPT-5.1 by 2.695.698.2+2.6
- HarmBenchtie98.197.6tie
- AIR-Bench 2024tie85.986.2tie
- anthropic_red_teamtie99.899.5tie
- simple_safety_teststie10099.8tie
Questions people ask
Which is better, Claude 3.5 Sonnet or GPT-5.1?
GPT-5.1 wins all three areas where both have results: coding, reasoning and images and charts. Claude 3.5 Sonnet wins none.
Which is better for coding?
GPT-5.1. It wins 3 of the 3 coding tests both models report; Claude 3.5 Sonnet wins none.
Which is cheaper?
Claude 3.5 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens; GPT-5.1 costs $1.25 and $10.00. That makes GPT-5.1 about 38% cheaper for the same work.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.