VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

DeepSeek-V4-Pro vs MiMo V2.5 Pro

38 SHARED BENCHMARKS

Across 38 shared benchmarks, DeepSeek-V4-Pro scores higher on 30 and MiMo V2.5 Pro on 8. The widest gap is AA ApexAgents, where DeepSeek-V4-Pro scores 24.3 against 2.4. MiMo V2.5 Pro is the cheaper of the two on tracked API pricing ($0.43 against $1.32 per million input tokens).

DEEPSEEKVSXIAOMI38 SHARED308 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerDeepSeekXiaomi
Released
Price per 1M tokens input / output$1.32 / $3.96$0.43 / $0.87
Cost of 1M in + 1M out$5.28$1.30 4.1× less
Head-to-head of 38 shared benchmarks30 wins8 wins
Scores tracked independently verified264 27101 4

Release dates per Artificial Analysis. Prices: DeepSeek's own price page for DeepSeek-V4-Pro; Artificial Analysis for MiMo V2.5 Pro. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaDeepSeek-V4-ProWINSMiMo V2.5 Pro
Reasoning30DeepSeek-V4-Pro leads 3 of 3 · widest: Humanity's Last Exam 37.5 vs 14.8
Coding52DeepSeek-V4-Pro leads 5 of 7 · widest: LiveCodeBench v6 92.5 vs 39.6
Agentic50DeepSeek-V4-Pro leads 5 of 5 · widest: AA ApexAgents 24.3 vs 2.4
Factuality21DeepSeek-V4-Pro leads 2 of 3 · widest: SimpleQA Verified 46.2 vs 16.1
Instruction Following10DeepSeek-V4-Pro leads 1 of 1 · widest: IFBench 76.5 vs 42.7
Long Context10DeepSeek-V4-Pro leads 1 of 1 · widest: AA-LCR 74.7 vs 41.7
Math20DeepSeek-V4-Pro leads 2 of 2 · widest: HMMT Feb. 2026 95.2 vs 82.6

Biggest gaps

DeepSeek-V4-Pro pulls furthest ahead on

  1. AA ApexAgents24.3 vs 2.4
  2. τ-Bench V3 · Banking30.1 vs 9.9
  3. SimpleQA Verified46.2 vs 16.1

MiMo V2.5 Pro pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination10.8 vs 5.9

Every shared benchmark38 · grouped by area

Reasoning 3

Agentic 5

GDPVal32.230.4
Terminal-Bench 4.014.60

Instruction Following 1

Long Context 1

AA-LCR74.741.7

Math 2

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (deepseek-official, direct), otherwise the lowest tracked offer.

Compare DeepSeek-V4-Pro withALL PAIRINGS →

Compare MiMo V2.5 Pro withALL PAIRINGS →