Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 5.5 vs GLM-5.2

27 SHARED BENCHMARKS

Across 27 shared benchmarks, Claude Opus 5.5 scores higher on 25 and GLM-5.2 on 2. The widest gap is ARC-AGI-2, where Claude Opus 5.5 scores 91.7 against 22.8. GLM-5.2 is the cheaper of the two on tracked API pricing ($1.40 against $4.00 per million input tokens).

ANTHROPICVSZ.AI27 SHARED25–2 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicZ.ai
Released
Price per 1M tokens input / output$4.00 / $20.00$1.40 / $4.40
Cost of 1M in + 1M out$24.00$5.80 4.1× less
Head-to-head of 27 shared benchmarks25 wins2 wins
Scores tracked independently verified98 25 ◆123 21 ◆

Release dates: the vendor's own announcement for Claude Opus 5.5; Artificial Analysis for GLM-5.2. Prices: Anthropic's own price page for Claude Opus 5.5; Z.ai's own price page for GLM-5.2. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 5.5WINSGLM-5.2
Reasoning50Claude Opus 5.5 leads 5 of 5 · widest: ARC-AGI-2 91.7 vs 22.8
Coding50Claude Opus 5.5 leads 5 of 5 · widest: SWE-bench Pro 89.9 vs 62.1
Agentic21Claude Opus 5.5 leads 2 of 3 · widest: GDPVal 67.3 vs 42.9
Factuality11Even, 1–1 of 2
Instruction Following10Claude Opus 5.5 leads 1 of 1 · widest: LiveBench · Instruction Following 65.7 vs 62.3
Long Context10Claude Opus 5.5 leads 1 of 1 · widest: AA-LCR 84.7 vs 78.3
Math10Claude Opus 5.5 leads 1 of 1 · widest: LiveBench · Mathematics 97.1 vs 89.8

Biggest gaps

Claude Opus 5.5 pulls furthest ahead on

  1. ARC-AGI-291.7 vs 22.8
  2. AA-Omniscience · Accuracy66.2 vs 24.3
  3. GDPVal67.3 vs 42.9

GLM-5.2 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination73.7 vs 41.4
  2. AA IT-Bench SRE42.7 vs 38.2

Every shared benchmark27 · grouped by area

Reasoning 5

BenchmarkClaude Opus 5.5MARGINGLM-5.2
ARC-AGI-291.7 ◆◆ 22.8
CritPt31.720.9
LiveBench · Reasoning92.2 ◆◆ 78.6
SimpleBench88.4 ◆◆ 58.8

Coding 5

BenchmarkClaude Opus 5.5MARGINGLM-5.2
LMArena · WebDev1827 ◆◆ 1602
LiveBench · Coding89.3 ◆◆ 79.7
SciCode66.951.2

Agentic 3

BenchmarkClaude Opus 5.5MARGINGLM-5.2
GDPVal67.342.9
Terminal-Bench 4.059.61

Instruction Following 1

Long Context 1

BenchmarkClaude Opus 5.5MARGINGLM-5.2
AA-LCR84.778.3

Math 1

BenchmarkClaude Opus 5.5MARGINGLM-5.2
LiveBench · Mathematics97.1 ◆◆ 89.8

Other shared benchmarks 9

BenchmarkClaude Opus 5.5MARGINGLM-5.2
ARC-AGI-197.5 ◆◆ 77
livebench_data_analysis80.3 ◆◆ 73.7
livebench_language86.3 ◆◆ 76.2
OfficeQA Pro67.741.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, zai-official), otherwise the lowest tracked offer.

Compare Claude Opus 5.5 withALL PAIRINGS →

Compare GLM-5.2 withALL PAIRINGS →