VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Claude Opus 5.5 vs Claude Opus 5

29 SHARED BENCHMARKS

Across 29 shared benchmarks, Claude Opus 5.5 scores higher on 27 and Claude Opus 5 on 2. The widest gap is Terminal-Bench-Science 0.1, where Claude Opus 5.5 scores 58.7 against 29. Claude Opus 5.5 is the cheaper of the two on tracked API pricing ($4.00 against $5.00 per million input tokens).

ANTHROPICVSANTHROPIC29 SHARED272 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicAnthropic
Released
Price per 1M tokens input / output$4.00 / $20.00$5.00 / $25.00
Cost of 1M in + 1M out$24.00 1.3× less$30.00
Head-to-head of 29 shared benchmarks27 wins2 wins
Scores tracked independently verified37 1194 22

Release dates per the vendor's own announcement. Prices: Artificial Analysis for Claude Opus 5.5; Anthropic's own price page for Claude Opus 5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 5.5WINSClaude Opus 5
Reasoning10Claude Opus 5.5 leads 1 of 1 · widest: Humanity's Last Exam 67.7 vs 51.3
Coding20Claude Opus 5.5 leads 2 of 2 · widest: SWE-bench Pro 89.9 vs 79.2

Biggest gaps

Claude Opus 5.5 pulls furthest ahead on

  1. Humanity's Last Exam67.7 vs 51.3
  2. SWE-bench Pro89.9 vs 79.2
  3. SWE-bench Multilingual93.9 vs 89.5

Claude Opus 5 pulls furthest ahead on

No ratified-area lead of 3 points or more.

Every shared benchmark29 · grouped by area

Reasoning 1

Other shared benchmarks 26

AA Intelligence58 44.8
AECI169.4165.2
CoBench 2.155.853.2
DeepSWE 1.174.268.8
FrontierCode v1.1 (Main)54.448
GDPval-AA 2.118461708
GMMLU94.392.5
HealthBench (raw score)68.167.1
HealthBench Professional (raw score)77.173.4
MILU93.192.1
OfficeQA78.978.1
OfficeQA Pro67.766.9
OSWorld 2.0 (partial score)81.874
OSWorld 2.0 (partial)81.875.4
OSWorld 2.0 (strict pass rate)48.737.2
OSWorld 2.0 (strict)48.739.6
OSWorld 2.0 partial81.874
Terminal-Bench 4.066.452.3
Terminal-Bench-Science 0.158.729

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, anthropic-official), otherwise the lowest tracked offer.

Compare Claude Opus 5 withALL PAIRINGS →