Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 4.7 vs Claude Opus 5.5

32 SHARED BENCHMARKS

Across 32 shared benchmarks, Claude Opus 4.7 scores higher on 5 and Claude Opus 5.5 on 27. The widest gap is FrontierMath Tier 4, where Claude Opus 5.5 scores 95 against 31.7. Claude Opus 5.5 is the cheaper of the two on tracked API pricing ($4.00 against $5.00 per million input tokens).

ANTHROPICVSANTHROPIC32 SHARED5–27 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicAnthropic
Released
Price per 1M tokens input / output$5.00 / $25.00$4.00 / $20.00
Cost of 1M in + 1M out$30.00$24.00 1.3× less
Head-to-head of 32 shared benchmarks5 wins27 wins
Scores tracked independently verified167 39 ◆101 28 ◆

Release dates: Artificial Analysis for Claude Opus 4.7; the vendor's own announcement for Claude Opus 5.5. Prices: Artificial Analysis for Claude Opus 4.7; Anthropic's own price page for Claude Opus 5.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 4.7WINSClaude Opus 5.5
Reasoning05Claude Opus 5.5 leads 5 of 5 · widest: CritPt 12 vs 31.7
Coding06Claude Opus 5.5 leads 6 of 6 · widest: LiveBench · Agentic Coding 50.7 vs 71.7
Agentic11Even, 1–1 of 2
Factuality12Claude Opus 5.5 leads 2 of 3 · widest: SimpleQA Verified 51.7 vs 72.2
Instruction Following10Claude Opus 4.7 leads 1 of 1
Long Context01Claude Opus 5.5 leads 1 of 1 · widest: AA-LCR 78.7 vs 84.7
Math03Claude Opus 5.5 leads 3 of 3 · widest: FrontierMath Tier 4 31.7 vs 95
Multimodal01Claude Opus 5.5 leads 1 of 1 · widest: MMMU-Pro 78.8 vs 87.7

Biggest gaps

Claude Opus 4.7 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination57.7 vs 41.4
  2. AA IT-Bench SRE46.7 vs 38.2

Claude Opus 5.5 pulls furthest ahead on

  1. FrontierMath Tier 495 vs 31.7
  2. CritPt31.7 vs 12
  3. GDPVal67.3 vs 41.9

Every shared benchmark32 · grouped by area

Reasoning 5

ARC-AGI-275.8◆ 91.7
CritPt1231.7
LiveBench · Reasoning87.2 ◆◆ 92.2
SimpleBench61.7 ◆◆ 88.4

Coding 6

Agentic 2

Instruction Following 1

Long Context 1

Math 3

Multimodal 1

Other shared benchmarks 10

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, anthropic-official), otherwise the lowest tracked offer.

Compare Claude Opus 4.7 withALL PAIRINGS →

Compare Claude Opus 5.5 withALL PAIRINGS →