VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

MiMo V2.5 Pro vs Qwen3.7 Max

35 SHARED BENCHMARKS

Across 35 shared benchmarks, MiMo V2.5 Pro scores higher on 2 and Qwen3.7 Max on 33. The widest gap is AA-Omniscience · Non-hallucination, where Qwen3.7 Max scores 74.4 against 10.8. MiMo V2.5 Pro is the cheaper of the two on tracked API pricing ($0.43 against $2.50 per million input tokens).

XIAOMIVSALIBABA35 SHARED233 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerXiaomiAlibaba
Released
Price per 1M tokens input / output$0.43 / $0.87$2.50 / $7.50
Cost of 1M in + 1M out$1.30 7.7× less$10.00
Head-to-head of 35 shared benchmarks2 wins33 wins
Scores tracked independently verified101 486 13

Release dates per Artificial Analysis. Prices: Artificial Analysis for MiMo V2.5 Pro; Alibaba's own price page for Qwen3.7 Max. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaMiMo V2.5 ProWINSQwen3.7 Max
Reasoning03Qwen3.7 Max leads 3 of 3 · widest: Humanity's Last Exam 14.8 vs 40.5
Coding16Qwen3.7 Max leads 6 of 7 · widest: LiveCodeBench v6 39.6 vs 91.6
Agentic04Qwen3.7 Max leads 4 of 4 · widest: AA IT-Bench SRE 38.2 vs 42.5
Factuality03Qwen3.7 Max leads 3 of 3 · widest: AA-Omniscience · Non-hallucination 10.8 vs 74.4
Instruction Following01Qwen3.7 Max leads 1 of 1 · widest: IFBench 42.7 vs 80.5
Long Context01Qwen3.7 Max leads 1 of 1 · widest: AA-LCR 41.7 vs 79
Math02Qwen3.7 Max leads 2 of 2 · widest: HMMT Feb. 2026 82.6 vs 97.1

Biggest gaps

MiMo V2.5 Pro pulls furthest ahead on

No ratified-area lead of 3 points or more.

Qwen3.7 Max pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination74.4 vs 10.8
  2. SimpleQA Verified55.8 vs 16.1
  3. Humanity's Last Exam40.5 vs 14.8

Every shared benchmark35 · grouped by area

Reasoning 3

Agentic 4

GDPVal30.430.7
Terminal-Bench 4.001.5

Instruction Following 1

IFBench42.780.5

Long Context 1

AA-LCR41.779

Math 2

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, alibaba-official), otherwise the lowest tracked offer.

Compare MiMo V2.5 Pro withALL PAIRINGS →

Compare Qwen3.7 Max withALL PAIRINGS →