Claude Opus 5.5 vs GPT-5.6 Sol
14 SHARED BENCHMARKSAcross 14 shared benchmarks, Claude Opus 5.5 scores higher on 14 and GPT-5.6 Sol on 0. The widest gap is Program Bench, where Claude Opus 5.5 scores 91.2 against 25.
ANTHROPICVSOPENAI14 SHARED14–0 HEAD-TO-HEADNEWEST SCORE ADDED
At a glance
MakerAnthropicOpenAI
Released
Price per 1M tokens input / output$4.00 / $20.00$4.00 / $20.00
Cost of 1M in + 1M out$24.00$24.00
Head-to-head of 14 shared benchmarks14 wins0 wins
Scores tracked independently verified37 1 ◆220 27 ◆
Release dates: the vendor's own announcement for Claude Opus 5.5; Artificial Analysis for GPT-5.6 Sol. Prices: Artificial Analysis for Claude Opus 5.5; Artificial Analysis for GPT-5.6 Sol. ◆ = independently verified score.
Where each leadsby capability area · benchmark wins
Biggest gaps
Claude Opus 5.5 pulls furthest ahead on
- SWE-bench Pro89.9 vs 64.6
- Humanity's Last Exam67.7 vs 49.5
GPT-5.6 Sol pulls furthest ahead on
No ratified-area lead of 3 points or more.
Every shared benchmark14 · grouped by area
Reasoning 1
Coding 1
Other shared benchmarks 12
FrontierCode v1.1 (Main)54.447.5
GDPval-AA 2.118461588
Terminal-Bench 4.066.437.3
Terminal-Bench-Science 0.158.722.4
Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.