GPT-5.6 Sol vs Sonnet 5.5
Wins 3 of 6 areas
Facts · Math · Following instructions
Wins 3 of 6 areas
Coding · Reasoning · Images and charts
The two are evenly matched.GPT-5.6 Sol is better at math and facts; Sonnet 5.5 at coding and reasoning.
Both rank among the ten best models we track in math.
Scores updated · 27 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software03Sonnet 5.53 of 4 tests · 1 tie
- ReasoningHard problems that need careful thinking02Sonnet 5.52 of 3 tests · 1 tie
- Images and chartsUnderstanding pictures, charts and video01Sonnet 5.51 of 1 test
- MathCompetition and research-level math10GPT-5.6 Sol1 of 3 tests · 2 ties
- FactsGetting facts right instead of making them up10GPT-5.6 Sol1 of 1 test
- Following instructionsDoing exactly what it is asked10GPT-5.6 Sol1 of 1 test
Facts, images and charts and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Sonnet 5.5 costs 50% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Sonnet 5.5 pulls ahead
- Long, multi-file coding tasks in real codebasesSWE-bench Pro+16.7points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+15points ahead
- Common-sense trick questionsSimpleBench+11.1points ahead
Where GPT-5.6 Sol pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+23.2points ahead
- Follows detailed instructions exactlyLiveBench · Instruction Following+15.1points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+2.5points ahead
Every test, side by side
All 27 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingSonnet 5.5
- SWE-bench ProSonnet 5.5 by 16.764.681.3+16.7
- LMArena · WebDevSonnet 5.5 by 166 rating points16201786+166 rating
- LiveBench · CodingSonnet 5.5 by 7.583.991.4+7.5
- LiveBench · Agentic Codingtie56.256.3tie
ReasoningSonnet 5.5
- Humanity's Last ExamSonnet 5.5 by 1549.564.5+15
- SimpleBenchSonnet 5.5 by 11.164.875.9+11.1
- LiveBench · Reasoningtie91.791.6tie
Images and chartsSonnet 5.5
- LMArena · VisionSonnet 5.5 by 11 rating points12791290+11 rating
MathGPT-5.6 Sol
- FrontierMath Tier 4GPT-5.6 Sol by 2.58380.5+2.5
- FrontierMath Tiers 1-3 (v2)tie89.188.8tie
- LiveBench · Mathematicstie96.296.1tie
Following instructionsGPT-5.6 Sol
- LiveBench · Instruction FollowingGPT-5.6 Sol by 15.171.856.8+15.1
Other results14 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Program BenchSonnet 5.5 by 54.72579.7+54.7
- Terminal-Bench-Science 0.1Sonnet 5.5 by 37.522.459.9+37.5
- Terminal-Bench 4.0Sonnet 5.5 by 33.337.370.6+33.3
- GDPval-AA 2.1Sonnet 5.5 by 252 rating points15881840+252 rating
- livebench_data_analysisGPT-5.6 Sol by 20.379.859.5+20.3
- livebench_languageGPT-5.6 Sol by 9.787.778+9.7
- AA IntelligenceSonnet 5.5 by 94756+9
- HealthBench ProfessionalSonnet 5.5 by 8.760.569.2+8.7
- HealthBench Professional length-adjustedSonnet 5.5 by 8.760.569.2+8.7
- HealthBenchSonnet 5.5 by 8.45765.4+8.4
- Toolathlon VerifiedSonnet 5.5 by 2.974.977.8+2.9
- OfficeQA ProSonnet 5.5 by 2.463.265.6+2.4
- DeepSWE 1.1GPT-5.6 Sol by 27371+2
- HLE (with tools)tie64.564.5tie
Questions people ask
Which is better, GPT-5.6 Sol or Sonnet 5.5?
GPT-5.6 Sol and Sonnet 5.5 each win three of the six areas where both have results. GPT-5.6 Sol wins facts, math and following instructions; Sonnet 5.5 wins coding, reasoning and images and charts. Sonnet 5.5 costs 50% less.
Which is better for coding?
Sonnet 5.5. It wins 3 of the 4 coding tests both models report; GPT-5.6 Sol wins none, and 1 is a tie.
Which is cheaper?
GPT-5.6 Sol costs $4.00 per million input tokens and $20.00 per million output tokens; Sonnet 5.5 costs $2.00 and $10.00. That makes Sonnet 5.5 about 50% cheaper for the same work.
How do you compare the two?
We use the 27 benchmark tests both models have published scores on. The verdict counts the 13 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 14 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.