VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs MiMo V2.5 Pro

38 SHARED BENCHMARKS

Across 38 shared benchmarks, DeepSeek-V4-Flash scores higher on 28 and MiMo V2.5 Pro on 9, with 1 level. The widest gap is τ-Bench V3 · Banking, where DeepSeek-V4-Flash scores 39.4 against 9.9. MiMo V2.5 Pro is the cheaper of the two on tracked API pricing ($0.43 against $0.44 per million input tokens).

DEEPSEEKVSXIAOMI38 SHARED289 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerDeepSeekXiaomi
Released
Price per 1M tokens input / output$0.44 / $1.32$0.43 / $0.87
Cost of 1M in + 1M out$1.76$1.30 1.4× less
Head-to-head of 38 shared benchmarks28 wins9 wins
Scores tracked independently verified257 32101 4

Release dates per Artificial Analysis. Prices: Artificial Analysis for DeepSeek-V4-Flash; Artificial Analysis for MiMo V2.5 Pro. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaDeepSeek-V4-FlashWINSMiMo V2.5 Pro
Reasoning30DeepSeek-V4-Flash leads 3 of 3 · widest: Humanity's Last Exam 38.6 vs 14.8
Coding33Even, 3–3 of 7
Agentic31DeepSeek-V4-Flash leads 3 of 4 · widest: τ-Bench V3 · Banking 39.4 vs 9.9
Factuality21DeepSeek-V4-Flash leads 2 of 3 · widest: SimpleQA Verified 34.1 vs 16.1
Instruction Following10DeepSeek-V4-Flash leads 1 of 1 · widest: IFBench 79.2 vs 42.7
Long Context10DeepSeek-V4-Flash leads 1 of 1 · widest: AA-LCR 79.7 vs 41.7
Math20DeepSeek-V4-Flash leads 2 of 2 · widest: HMMT Feb. 2026 94.8 vs 82.6

Biggest gaps

DeepSeek-V4-Flash pulls furthest ahead on

  1. τ-Bench V3 · Banking39.4 vs 9.9
  2. Humanity's Last Exam38.6 vs 14.8
  3. LiveCodeBench v690.6 vs 39.6

MiMo V2.5 Pro pulls furthest ahead on

  1. AA IT-Bench SRE38.2 vs 31.5
  2. SWE-bench Pro57.2 vs 52.6
  3. LMArena · WebDev1437 vs 1431

Every shared benchmark38 · grouped by area

Reasoning 3

Agentic 4

GDPVal46.330.4
Terminal-Bench 4.012.10

Instruction Following 1

Long Context 1

Math 2

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare DeepSeek-V4-Flash withALL PAIRINGS →

Compare MiMo V2.5 Pro withALL PAIRINGS →