VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Claude Opus 4.8 vs MiMo V2.5 Pro

37 SHARED BENCHMARKS

Across 37 shared benchmarks, Claude Opus 4.8 scores higher on 37 and MiMo V2.5 Pro on 0. The widest gap is Program Bench, where Claude Opus 4.8 scores 71.9 against 12.5. MiMo V2.5 Pro is the cheaper of the two on tracked API pricing ($0.43 against $5.00 per million input tokens).

ANTHROPICVSXIAOMI37 SHARED370 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicXiaomi
Released
Price per 1M tokens input / output$5.00 / $25.00$0.43 / $0.87
Cost of 1M in + 1M out$30.00$1.30 23.1× less
Head-to-head of 37 shared benchmarks37 wins0 wins
Scores tracked independently verified201 32101 4

Release dates per Artificial Analysis. Prices: Artificial Analysis for Claude Opus 4.8; Artificial Analysis for MiMo V2.5 Pro. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Opus 4.8WINSMiMo V2.5 Pro
Reasoning30Claude Opus 4.8 leads 3 of 3 · widest: Humanity's Last Exam 48.7 vs 14.8
Coding60Claude Opus 4.8 leads 6 of 6 · widest: Terminal-Bench Hard 58.3 vs 35.6
Agentic40Claude Opus 4.8 leads 4 of 4 · widest: τ-Bench V3 · Banking 34.2 vs 9.9
Factuality30Claude Opus 4.8 leads 3 of 3 · widest: AA-Omniscience · Non-hallucination 60.7 vs 10.8
Instruction Following10Claude Opus 4.8 leads 1 of 1 · widest: IFBench 62.2 vs 42.7
Long Context10Claude Opus 4.8 leads 1 of 1 · widest: AA-LCR 77.7 vs 41.7
Math20Claude Opus 4.8 leads 2 of 2 · widest: HMMT Feb. 2026 95.5 vs 82.6
Multimodal20Claude Opus 4.8 leads 2 of 2 · widest: CharXiv (RQ) 89.9 vs 81

Biggest gaps

Claude Opus 4.8 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination60.7 vs 10.8
  2. τ-Bench V3 · Banking34.2 vs 9.9
  3. SimpleQA Verified53 vs 16.1

MiMo V2.5 Pro pulls furthest ahead on

No ratified-area lead of 3 points or more.

Every shared benchmark37 · grouped by area

Reasoning 3

Agentic 4

GDPVal46.930.4
Terminal-Bench 4.021.70

Instruction Following 1

Long Context 1

AA-LCR77.741.7

Math 2

AIME 2026100 93.6
HMMT Feb. 202695.5 82.6

Multimodal 2

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare Claude Opus 4.8 withALL PAIRINGS →

Compare MiMo V2.5 Pro withALL PAIRINGS →