VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

MiMo V2.5 Pro vs Qwen3.5 397B A17B

39 SHARED BENCHMARKS

Across 39 shared benchmarks, MiMo V2.5 Pro scores higher on 15 and Qwen3.5 397B A17B on 23, with 1 level. The widest gap is AA ApexAgents, where Qwen3.5 397B A17B scores 15.3 against 2.4. MiMo V2.5 Pro is the cheaper of the two on tracked API pricing ($0.43 against $0.60 per million input tokens).

XIAOMIVSALIBABA39 SHARED1523 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerXiaomiAlibaba
Released
Price per 1M tokens input / output$0.43 / $0.87$0.60 / $3.60
Cost of 1M in + 1M out$1.30 3.2× less$4.20
Head-to-head of 39 shared benchmarks15 wins23 wins
Scores tracked independently verified101 4162 11

Release dates per Artificial Analysis. Prices: Artificial Analysis for MiMo V2.5 Pro; Artificial Analysis for Qwen3.5 397B A17B. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaMiMo V2.5 ProWINSQwen3.5 397B A17B
Reasoning03Qwen3.5 397B A17B leads 3 of 3 · widest: Humanity's Last Exam 14.8 vs 29
Coding52MiMo V2.5 Pro leads 5 of 7 · widest: Terminal-Bench 2.1 65.2 vs 51.3
Agentic22Even, 2–2 of 5
Factuality03Qwen3.5 397B A17B leads 3 of 3 · widest: SimpleQA Verified 16.1 vs 26
Instruction Following01Qwen3.5 397B A17B leads 1 of 1 · widest: IFBench 42.7 vs 78.8
Long Context01Qwen3.5 397B A17B leads 1 of 1 · widest: AA-LCR 41.7 vs 77.3
Math11Even, 1–1 of 2
Multimodal11Even, 1–1 of 2

Biggest gaps

MiMo V2.5 Pro pulls furthest ahead on

  1. GDPVal30.4 vs 14
  2. Terminal-Bench 2.165.2 vs 51.3
  3. SciCode50.6 vs 44.8

Qwen3.5 397B A17B pulls furthest ahead on

  1. AA ApexAgents15.3 vs 2.4
  2. LiveCodeBench v683.6 vs 39.6
  3. Humanity's Last Exam29 vs 14.8

Every shared benchmark39 · grouped by area

Reasoning 3

Agentic 5

GDPVal30.414
Terminal-Bench 4.000

Instruction Following 1

Long Context 1

Math 2

Multimodal 2

LMArena · Vision1247 1265
MMMU-Pro77.977.3

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare MiMo V2.5 Pro withALL PAIRINGS →

Compare Qwen3.5 397B A17B withALL PAIRINGS →