Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

MiMo V2.5 vs Qwen3.5 397B A17B

31 SHARED BENCHMARKS

Across 31 shared benchmarks, MiMo V2.5 scores higher on 14 and Qwen3.5 397B A17B on 16, with 1 level. The widest gap is AA-Omniscience · Non-hallucination, where MiMo V2.5 scores 68.1 against 11.1. MiMo V2.5 is the cheaper of the two on tracked API pricing ($0.14 against $0.60 per million input tokens).

XIAOMIVSALIBABA31 SHARED14–16 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerXiaomiAlibaba
Released
Price per 1M tokens input / output$0.14 / $0.28$0.60 / $3.60
Cost of 1M in + 1M out$0.42 10.0× less$4.20
Head-to-head of 31 shared benchmarks14 wins16 wins
Scores tracked independently verified51 3 ◆162 11 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for MiMo V2.5; Alibaba's own price page for Qwen3.5 397B A17B. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaMiMo V2.5WINSQwen3.5 397B A17B
Reasoning12Qwen3.5 397B A17B leads 2 of 3 · widest: GPQA Diamond 84.9 vs 89.3
Coding42MiMo V2.5 leads 4 of 6 · widest: Terminal-Bench 2.1 63.7 vs 51.3
Agentic11Even, 1–1 of 3
Factuality12Qwen3.5 397B A17B leads 2 of 3 · widest: AA-Omniscience · Accuracy 16.8 vs 30.8
Instruction Following01Qwen3.5 397B A17B leads 1 of 1 · widest: IFBench 67.1 vs 78.8
Long Context01Qwen3.5 397B A17B leads 1 of 1 · widest: AA-LCR 73 vs 77.3
Math11Even, 1–1 of 2
Multimodal02Qwen3.5 397B A17B leads 2 of 2 · widest: LMArena · Vision 1247 vs 1265

Biggest gaps

MiMo V2.5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination68.1 vs 11.1
  2. GDPVal24.3 vs 14
  3. Terminal-Bench 2.163.7 vs 51.3

Qwen3.5 397B A17B pulls furthest ahead on

  1. AA-Omniscience · Accuracy30.8 vs 16.8
  2. SimpleQA Verified26 vs 16.1
  3. τ-Bench V3 · Banking13.4 vs 8.7

Every shared benchmark31 · grouped by area

Reasoning 3

Agentic 3

GDPVal24.314
Terminal-Bench 4.000

Instruction Following 1

IFBench67.178.8

Long Context 1

Math 2

AIME 202693.691.3
HMMT Feb. 202682.6◆ 87.9

Multimodal 2

LMArena · Vision1247 ◆◆ 1265
MMMU-Pro75.477.3

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, alibaba-official), otherwise the lowest tracked offer.

Compare Qwen3.5 397B A17B withALL PAIRINGS →