Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 5.5 vs Qwen3.8 Max Preview

27 SHARED BENCHMARKS

Across 27 shared benchmarks, Claude Opus 5.5 scores higher on 24 and Qwen3.8 Max Preview on 3. The widest gap is AA-Omniscience, where Claude Opus 5.5 scores 46.4 against 12. Qwen3.8 Max Preview is the cheaper of the two on tracked API pricing ($2.00 against $4.00 per million input tokens).

ANTHROPICVSALIBABA27 SHARED24–3 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicAlibaba
Released
Price per 1M tokens input / output$4.00 / $20.00$2.00 / $6.00
Cost of 1M in + 1M out$24.00$8.00 3.0× less
Head-to-head of 27 shared benchmarks24 wins3 wins
Scores tracked independently verified101 28 ◆87 19 ◆

Release dates: the vendor's own announcement for Claude Opus 5.5; Artificial Analysis for Qwen3.8 Max Preview. Prices: Artificial Analysis for Claude Opus 5.5; Artificial Analysis for Qwen3.8 Max Preview. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 5.5WINSQwen3.8 Max Preview
Reasoning30Claude Opus 5.5 leads 3 of 3 · widest: CritPt 31.7 vs 17.7
Coding50Claude Opus 5.5 leads 5 of 5 · widest: SWE-bench Pro 89.9 vs 67.7
Agentic21Claude Opus 5.5 leads 2 of 3 · widest: Terminal-Bench 4.0 59.6 vs 38.9
Factuality21Claude Opus 5.5 leads 2 of 3 · widest: AA-Omniscience · Accuracy 66.2 vs 31.7
Instruction Following01Qwen3.8 Max Preview leads 1 of 1 · widest: LiveBench · Instruction Following 65.7 vs 74.1
Long Context10Claude Opus 5.5 leads 1 of 1 · widest: AA-LCR 84.7 vs 80.3
Math30Claude Opus 5.5 leads 3 of 3 · widest: FrontierMath Tier 4 95 vs 46.3
Multimodal10Claude Opus 5.5 leads 1 of 1 · widest: MMMU-Pro 87.7 vs 82.8

Biggest gaps

Claude Opus 5.5 pulls furthest ahead on

  1. AA-Omniscience · Accuracy66.2 vs 31.7
  2. FrontierMath Tier 495 vs 46.3
  3. CritPt31.7 vs 17.7

Qwen3.8 Max Preview pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination71.2 vs 41.4
  2. LiveBench · Instruction Following74.1 vs 65.7

Every shared benchmark27 · grouped by area

Reasoning 3

Coding 5

LMArena · WebDev1820 ◆◆ 1671
LiveBench · Coding89.3 ◆◆ 72.9
SciCode66.952.1

Agentic 3

GDPVal67.358.2
Terminal-Bench 4.059.638.9

Instruction Following 1

Long Context 1

Multimodal 1

Other shared benchmarks 7

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare Claude Opus 5.5 withALL PAIRINGS →

Compare Qwen3.8 Max Preview withALL PAIRINGS →