Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Kimi K2.5 vs MiMo V2.5

31 SHARED BENCHMARKS

Across 31 shared benchmarks, Kimi K2.5 scores higher on 17 and MiMo V2.5 on 14. The widest gap is SimpleQA Verified, where Kimi K2.5 scores 34.3 against 16.1. MiMo V2.5 is the cheaper of the two on tracked API pricing ($0.14 against $0.60 per million input tokens).

MOONSHOTVSXIAOMI31 SHARED17–14 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerMoonshotXiaomi
Released
Price per 1M tokens input / output$0.60 / $3.00$0.14 / $0.28
Cost of 1M in + 1M out$3.60$0.42 8.6× less
Head-to-head of 31 shared benchmarks17 wins14 wins
Scores tracked independently verified172 26 ◆51 3 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for Kimi K2.5; Artificial Analysis for MiMo V2.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaKimi K2.5WINSMiMo V2.5
Reasoning21Kimi K2.5 leads 2 of 3 · widest: Humanity's Last Exam 30.7 vs 27.2
Coding24MiMo V2.5 leads 4 of 6 · widest: Terminal-Bench 2.1 45.7 vs 63.7
Agentic11Even, 1–1 of 2
Factuality21Kimi K2.5 leads 2 of 3 · widest: SimpleQA Verified 34.3 vs 16.1
Instruction Following10Kimi K2.5 leads 1 of 1 · widest: IFBench 70.2 vs 67.1
Long Context10Kimi K2.5 leads 1 of 1 · widest: AA-LCR 78 vs 73
Math20Kimi K2.5 leads 2 of 2 · widest: HMMT Feb. 2026 87.1 vs 82.6
Multimodal13MiMo V2.5 leads 3 of 4 · widest: CharXiv (RQ) 77.5 vs 81

Biggest gaps

Kimi K2.5 pulls furthest ahead on

  1. SimpleQA Verified34.3 vs 16.1
  2. AA-Omniscience · Accuracy35.2 vs 16.8
  3. τ-Bench V3 · Banking14.2 vs 8.7

MiMo V2.5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination68.1 vs 34.3
  2. GDPVal24.3 vs 16.2
  3. Terminal-Bench 2.163.7 vs 45.7

Every shared benchmark31 · grouped by area

Reasoning 3

Coding 6

Agentic 2

Instruction Following 1

BenchmarkKimi K2.5MARGINMiMo V2.5
IFBench70.267.1

Long Context 1

BenchmarkKimi K2.5MARGINMiMo V2.5
AA-LCR7873

Math 2

BenchmarkKimi K2.5MARGINMiMo V2.5
AIME 202695.8 ◆93.6
HMMT Feb. 202687.1 ◆82.6

Multimodal 4

BenchmarkKimi K2.5MARGINMiMo V2.5
LMArena · Vision1267 ◆◆ 1247
MMMU-Pro75.475.4
Video-MME87.487.7

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare Kimi K2.5 withALL PAIRINGS →