Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 5.5 vs GPT-6 Sol

26 SHARED BENCHMARKS

Across 26 shared benchmarks, Claude Opus 5.5 scores higher on 23 and GPT-6 Sol on 3. The widest gap is AA-Omniscience, where Claude Opus 5.5 scores 46.4 against 27.1. GPT-6 Sol is the cheaper of the two on tracked API pricing ($2.00 against $4.00 per million input tokens).

ANTHROPICVSOPENAI26 SHARED23–3 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicOpenAI
Released
Price per 1M tokens input / output$4.00 / $20.00$2.00 / $10.00
Cost of 1M in + 1M out$24.00$12.00 2.0× less
Head-to-head of 26 shared benchmarks23 wins3 wins
Scores tracked independently verified98 25 ◆107 37 ◆

Release dates: the vendor's own announcement for Claude Opus 5.5; Artificial Analysis for GPT-6 Sol. Prices: Artificial Analysis for Claude Opus 5.5; OpenAI's own price page for GPT-6 Sol. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 5.5WINSGPT-6 Sol
Reasoning50Claude Opus 5.5 leads 5 of 5 · widest: Humanity's Last Exam 61.4 vs 47.9
Coding40Claude Opus 5.5 leads 4 of 4 · widest: LiveBench · Agentic Coding 71.7 vs 52.9
Agentic21Claude Opus 5.5 leads 2 of 3 · widest: GDPVal 67.3 vs 49.3
Factuality20Claude Opus 5.5 leads 2 of 2 · widest: AA-Omniscience · Accuracy 66.2 vs 54.5
Instruction Following01GPT-6 Sol leads 1 of 1
Long Context10Claude Opus 5.5 leads 1 of 1
Math10Claude Opus 5.5 leads 1 of 1
Multimodal10Claude Opus 5.5 leads 1 of 1 · widest: MMMU-Pro 87.7 vs 83.3

Biggest gaps

Claude Opus 5.5 pulls furthest ahead on

  1. GDPVal67.3 vs 49.3
  2. Terminal-Bench 4.059.6 vs 43.9
  3. LiveBench · Agentic Coding71.7 vs 52.9

GPT-6 Sol pulls furthest ahead on

  1. AA IT-Bench SRE49.4 vs 38.2

Every shared benchmark26 · grouped by area

Reasoning 5

ARC-AGI-291.7 ◆◆ 89.6
CritPt31.730.9
LiveBench · Reasoning92.2 ◆◆ 88.7
SimpleBench88.4 ◆◆ 73.1

Coding 4

LMArena · WebDev1827 ◆◆ 1681
LiveBench · Coding89.3 ◆◆ 81.8
SciCode66.957.6

Agentic 3

GDPVal67.349.3
Terminal-Bench 4.059.643.9

Instruction Following 1

Long Context 1

AA-LCR84.783.7

Math 1

Multimodal 1

MMMU-Pro87.783.3

Other shared benchmarks 8

ARC-AGI-197.5 ◆◆ 95.5
DeepSWE 1.174.268.8
GDPval-AA 2.118461487
livebench_data_analysis80.3 ◆◆ 81.2
livebench_language86.3 ◆◆ 85.3
OSWorld 2.081.864.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, openai-official), otherwise the lowest tracked offer.

Compare Claude Opus 5.5 withALL PAIRINGS →