GPT-5.6 Sol wins more areas, narrowly.
Scores updated · 12 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Agents and facts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
GPT-5.6 Sol costs 89% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-5.6 Sol pulls ahead
- Abstract visual puzzles that people can solveARC-AGI-2+8.3points ahead
- Short factual questions, answered correctlySimpleQA Verified+5.2points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+4.9points ahead
Where GPT-5.5 Pro pulls ahead
- Common-sense trick questionsSimpleBench+12.1points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+7.7points ahead
Every test, side by side
All 12 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningEven
- SimpleBenchGPT-5.5 Pro by 12.176.964.8+12.1
- ARC-AGI-2GPT-5.6 Sol by 8.384.292.5+8.3
- Humanity's Last ExamGPT-5.5 Pro by 7.757.249.5+7.7
- CritPtGPT-5.6 Sol by 1.730.632.3+1.7
MathGPT-5.6 Sol
- FrontierMath Tier 4GPT-5.6 Sol by 4.97883+4.9
- FrontierMath Tiers 1-3 (v2)GPT-5.6 Sol by 1.487.789.1+1.4
Other results4 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- FrontierMath (overall)GPT-5.6 Sol by 49.439.689+49.4
- Hard-negative protein binding predictionGPT-5.6 Sol by 7.607.6+7.6
- DNA sequence design for transcription factor bindingGPT-5.5 Pro by 2.816.513.7+2.8
- ARC-AGI-1GPT-5.6 Sol by 1.59596.5+1.5
Questions people ask
Which is better, GPT-5.5 Pro or GPT-5.6 Sol?
GPT-5.6 Sol wins two of the four areas where both have results: facts and math. GPT-5.5 Pro wins none. They are level on agents and reasoning.
Which is cheaper?
GPT-5.5 Pro costs $30.00 per million input tokens and $180.00 per million output tokens; GPT-5.6 Sol costs $4.00 and $20.00. That makes GPT-5.6 Sol about 89% cheaper for the same work.
How do you compare the two?
We use the 12 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 4 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.