Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 5.5 vs Gemini 3.1 Pro

30 SHARED BENCHMARKS

Across 30 shared benchmarks, Claude Opus 5.5 scores higher on 26 and Gemini 3.1 Pro on 4. The widest gap is DeepSWE 1.1, where Claude Opus 5.5 scores 74.2 against 12. Gemini 3.1 Pro is the cheaper of the two on tracked API pricing ($2.00 against $4.00 per million input tokens).

ANTHROPICVSGOOGLE30 SHARED26–4 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicGoogle
Released
Price per 1M tokens input / output$4.00 / $20.00$2.00 / $12.00
Cost of 1M in + 1M out$24.00$14.00 1.7× less
Head-to-head of 30 shared benchmarks26 wins4 wins
Scores tracked independently verified98 25 ◆215 30 ◆

Release dates: the vendor's own announcement for Claude Opus 5.5; Artificial Analysis for Gemini 3.1 Pro. Prices: Anthropic's own price page for Claude Opus 5.5; Google's own price page for Gemini 3.1 Pro. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 5.5WINSGemini 3.1 Pro
Reasoning50Claude Opus 5.5 leads 5 of 5 · widest: CritPt 31.7 vs 17.7
Coding60Claude Opus 5.5 leads 6 of 6 · widest: SWE-bench Pro 89.9 vs 54.2
Agentic30Claude Opus 5.5 leads 3 of 3 · widest: GDPVal 67.3 vs 13.8
Factuality11Even, 1–1 of 2
Instruction Following01Gemini 3.1 Pro leads 1 of 1 · widest: LiveBench · Instruction Following 65.7 vs 79.1
Long Context10Claude Opus 5.5 leads 1 of 1
Math10Claude Opus 5.5 leads 1 of 1 · widest: LiveBench · Mathematics 97.1 vs 91
Multimodal10Claude Opus 5.5 leads 1 of 1 · widest: MMMU-Pro 87.7 vs 82.4

Biggest gaps

Claude Opus 5.5 pulls furthest ahead on

  1. GDPVal67.3 vs 13.8
  2. CritPt31.7 vs 17.7
  3. SWE-bench Pro89.9 vs 54.2

Gemini 3.1 Pro pulls furthest ahead on

  1. LiveBench · Instruction Following79.1 vs 65.7
  2. AA-Omniscience · Non-hallucination49.1 vs 41.4

Every shared benchmark30 · grouped by area

Reasoning 5

ARC-AGI-291.7 ◆77.1
CritPt31.717.7
SimpleBench88.4 ◆◆ 79.6

Coding 6

Agentic 3

GDPVal67.313.8
Terminal-Bench 4.059.64

Instruction Following 1

Long Context 1

Multimodal 1

Other shared benchmarks 10

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, google-official), otherwise the lowest tracked offer.

Compare Claude Opus 5.5 withALL PAIRINGS →

Compare Gemini 3.1 Pro withALL PAIRINGS →