Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Kimi K2.6 vs MiMo V2.5

33 SHARED BENCHMARKS

Across 33 shared benchmarks, Kimi K2.6 scores higher on 31 and MiMo V2.5 on 2. The widest gap is τ-Bench V3 · Banking, where Kimi K2.6 scores 23.3 against 8.7. MiMo V2.5 is the cheaper of the two on tracked API pricing ($0.14 against $0.95 per million input tokens).

MOONSHOTVSXIAOMI33 SHARED31–2 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerMoonshotXiaomi
Released
Price per 1M tokens input / output$0.95 / $4.00$0.14 / $0.28
Cost of 1M in + 1M out$4.95$0.42 11.8× less
Head-to-head of 33 shared benchmarks31 wins2 wins
Scores tracked independently verified181 26 ◆51 3 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for Kimi K2.6; Artificial Analysis for MiMo V2.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaKimi K2.6WINSMiMo V2.5
Reasoning30Kimi K2.6 leads 3 of 3 · widest: CritPt 8 vs 3.7
Coding60Kimi K2.6 leads 6 of 6 · widest: SciCode 51.5 vs 43.9
Agentic30Kimi K2.6 leads 3 of 3 · widest: τ-Bench V3 · Banking 23.3 vs 8.7
Factuality21Kimi K2.6 leads 2 of 3 · widest: SimpleQA Verified 34.9 vs 16.1
Instruction Following10Kimi K2.6 leads 1 of 1 · widest: IFBench 76 vs 67.1
Long Context10Kimi K2.6 leads 1 of 1 · widest: AA-LCR 81 vs 73
Math20Kimi K2.6 leads 2 of 2 · widest: HMMT Feb. 2026 92.7 vs 82.6
Multimodal30Kimi K2.6 leads 3 of 3 · widest: CharXiv (RQ) 86.7 vs 81

Biggest gaps

Kimi K2.6 pulls furthest ahead on

  1. τ-Bench V3 · Banking23.3 vs 8.7
  2. SimpleQA Verified34.9 vs 16.1
  3. CritPt8 vs 3.7

MiMo V2.5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination68.1 vs 59.5

Every shared benchmark33 · grouped by area

Reasoning 3

Coding 6

Agentic 3

BenchmarkKimi K2.6MARGINMiMo V2.5
GDPVal26.324.3
Terminal-Bench 4.00.50

Instruction Following 1

BenchmarkKimi K2.6MARGINMiMo V2.5
IFBench7667.1

Long Context 1

BenchmarkKimi K2.6MARGINMiMo V2.5
AA-LCR8173

Math 2

BenchmarkKimi K2.6MARGINMiMo V2.5
AIME 202696.493.6

Multimodal 3

BenchmarkKimi K2.6MARGINMiMo V2.5
LMArena · Vision1280 ◆◆ 1247
MMMU-Pro79.475.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare Kimi K2.6 withALL PAIRINGS →