Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Opus 5.5 vs Kimi K3

27 SHARED BENCHMARKS

Across 27 shared benchmarks, Claude Opus 5.5 scores higher on 23 and Kimi K3 on 4. The widest gap is Terminal-Bench 4.0, where Claude Opus 5.5 scores 59.6 against 12.6. Kimi K3 is the cheaper of the two on tracked API pricing ($3.00 against $4.00 per million input tokens).

ANTHROPICVSMOONSHOT27 SHARED23–4 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicMoonshot
Released
Price per 1M tokens input / output$4.00 / $20.00$3.00 / $15.00
Cost of 1M in + 1M out$24.00$18.00 1.3× less
Head-to-head of 27 shared benchmarks23 wins4 wins
Scores tracked independently verified98 25 ◆131 26 ◆

Release dates: the vendor's own announcement for Claude Opus 5.5; Artificial Analysis for Kimi K3. Prices: Anthropic's own price page for Claude Opus 5.5; Artificial Analysis for Kimi K3. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 5.5WINSKimi K3
Reasoning50Claude Opus 5.5 leads 5 of 5 · widest: ARC-AGI-2 91.7 vs 60.4
Coding40Claude Opus 5.5 leads 4 of 4 · widest: LiveBench · Agentic Coding 71.7 vs 62.2
Agentic21Claude Opus 5.5 leads 2 of 3 · widest: Terminal-Bench 4.0 59.6 vs 12.6
Factuality11Even, 1–1 of 2
Instruction Following01Kimi K3 leads 1 of 1 · widest: LiveBench · Instruction Following 65.7 vs 71.4
Long Context01Kimi K3 leads 1 of 1 · widest: AA-LCR 84.7 vs 88.7
Math10Claude Opus 5.5 leads 1 of 1 · widest: LiveBench · Mathematics 97.1 vs 84.4
Multimodal10Claude Opus 5.5 leads 1 of 1 · widest: MMMU-Pro 87.7 vs 80.5

Biggest gaps

Claude Opus 5.5 pulls furthest ahead on

  1. Terminal-Bench 4.059.6 vs 12.6
  2. ARC-AGI-291.7 vs 60.4
  3. SimpleBench88.4 vs 60.7

Kimi K3 pulls furthest ahead on

  1. AA IT-Bench SRE47.7 vs 38.2
  2. AA-Omniscience · Non-hallucination46.8 vs 41.4
  3. LiveBench · Instruction Following71.4 vs 65.7

Every shared benchmark27 · grouped by area

Reasoning 5

BenchmarkClaude Opus 5.5MARGINKimi K3
ARC-AGI-291.7 ◆◆ 60.4
CritPt31.723.4
LiveBench · Reasoning92.2 ◆◆ 90.7
SimpleBench88.4 ◆◆ 60.7

Coding 4

BenchmarkClaude Opus 5.5MARGINKimi K3
LMArena · WebDev1827 ◆◆ 1660
LiveBench · Coding89.3 ◆◆ 81.5
SciCode66.959.5

Agentic 3

BenchmarkClaude Opus 5.5MARGINKimi K3
GDPVal67.351.2
Terminal-Bench 4.059.612.6

Instruction Following 1

Long Context 1

BenchmarkClaude Opus 5.5MARGINKimi K3
AA-LCR84.788.7

Math 1

BenchmarkClaude Opus 5.5MARGINKimi K3
LiveBench · Mathematics97.1 ◆◆ 84.4

Multimodal 1

BenchmarkClaude Opus 5.5MARGINKimi K3
MMMU-Pro87.780.5

Other shared benchmarks 9

BenchmarkClaude Opus 5.5MARGINKimi K3
ARC-AGI-197.5 ◆◆ 94.5
livebench_data_analysis80.3 ◆◆ 78.7
livebench_language86.3 ◆◆ 85.5
OfficeQA Pro67.763.3

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.

Compare Claude Opus 5.5 withALL PAIRINGS →

Compare Kimi K3 withALL PAIRINGS →