VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

MiMo V2.5 Pro vs Qwen3.6 Plus

35 SHARED BENCHMARKS

Across 35 shared benchmarks, MiMo V2.5 Pro scores higher on 13 and Qwen3.6 Plus on 22. The widest gap is AA-Omniscience · Non-hallucination, where Qwen3.6 Plus scores 65.4 against 10.8. MiMo V2.5 Pro is the cheaper of the two on tracked API pricing ($0.43 against $0.50 per million input tokens).

XIAOMIVSALIBABA35 SHARED1322 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerXiaomiAlibaba
Released
Price per 1M tokens input / output$0.43 / $0.87$0.50 / $3.00
Cost of 1M in + 1M out$1.30 2.7× less$3.50
Head-to-head of 35 shared benchmarks13 wins22 wins
Scores tracked independently verified101 4102 12

Release dates per Artificial Analysis. Prices: Artificial Analysis for MiMo V2.5 Pro; Alibaba's own price page for Qwen3.6 Plus. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaMiMo V2.5 ProWINSQwen3.6 Plus
Reasoning03Qwen3.6 Plus leads 3 of 3 · widest: Humanity's Last Exam 14.8 vs 27.8
Coding43MiMo V2.5 Pro leads 4 of 7 · widest: SciCode 50.6 vs 40.7
Agentic11Even, 1–1 of 2
Factuality12Qwen3.6 Plus leads 2 of 3 · widest: AA-Omniscience · Non-hallucination 10.8 vs 65.4
Instruction Following01Qwen3.6 Plus leads 1 of 1 · widest: IFBench 42.7 vs 75.2
Long Context01Qwen3.6 Plus leads 1 of 1 · widest: AA-LCR 41.7 vs 78.3
Math02Qwen3.6 Plus leads 2 of 2 · widest: HMMT Feb. 2026 82.6 vs 87.8
Multimodal12Qwen3.6 Plus leads 2 of 3

Biggest gaps

MiMo V2.5 Pro pulls furthest ahead on

  1. GDPVal30.4 vs 23.8
  2. SciCode50.6 vs 40.7
  3. Terminal-Bench 2.165.2 vs 61.4

Qwen3.6 Plus pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination65.4 vs 10.8
  2. SimpleQA Verified44.1 vs 16.1
  3. LiveCodeBench v687.1 vs 39.6

Every shared benchmark35 · grouped by area

Reasoning 3

Agentic 2

Instruction Following 1

IFBench42.775.2

Long Context 1

AA-LCR41.778.3

Math 2

Multimodal 3

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, alibaba-official), otherwise the lowest tracked offer.

Compare MiMo V2.5 Pro withALL PAIRINGS →

Compare Qwen3.6 Plus withALL PAIRINGS →