Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Gemini 3.1 Pro vs MiMo V2.5

32 SHARED BENCHMARKS

Across 32 shared benchmarks, Gemini 3.1 Pro scores higher on 26 and MiMo V2.5 on 6. The widest gap is CritPt, where Gemini 3.1 Pro scores 17.7 against 3.7. MiMo V2.5 is the cheaper of the two on tracked API pricing ($0.14 against $2.00 per million input tokens).

GOOGLEVSXIAOMI32 SHARED26–6 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerGoogleXiaomi
Released
Price per 1M tokens input / output$2.00 / $12.00$0.14 / $0.28
Cost of 1M in + 1M out$14.00$0.42 33.3× less
Head-to-head of 32 shared benchmarks26 wins6 wins
Scores tracked independently verified193 31 ◆51 3 ◆

Release dates per Artificial Analysis. Prices: Google's own price page for Gemini 3.1 Pro; Artificial Analysis for MiMo V2.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGemini 3.1 ProWINSMiMo V2.5
Reasoning30Gemini 3.1 Pro leads 3 of 3 · widest: CritPt 17.7 vs 3.7
Coding51Gemini 3.1 Pro leads 5 of 6 · widest: SciCode 58.7 vs 43.9
Agentic21Gemini 3.1 Pro leads 2 of 3 · widest: τ-Bench V3 · Banking 21.4 vs 8.7
Factuality21Gemini 3.1 Pro leads 2 of 3 · widest: SimpleQA Verified 73.5 vs 16.1
Instruction Following10Gemini 3.1 Pro leads 1 of 1 · widest: IFBench 77.1 vs 67.1
Long Context10Gemini 3.1 Pro leads 1 of 1 · widest: AA-LCR 82 vs 73
Math20Gemini 3.1 Pro leads 2 of 2 · widest: HMMT Feb. 2026 94.7 vs 82.6
Multimodal30Gemini 3.1 Pro leads 3 of 3 · widest: MMMU-Pro 82.4 vs 75.4

Biggest gaps

Gemini 3.1 Pro pulls furthest ahead on

  1. CritPt17.7 vs 3.7
  2. SimpleQA Verified73.5 vs 16.1
  3. AA-Omniscience · Accuracy54.9 vs 16.8

MiMo V2.5 pulls furthest ahead on

  1. GDPVal24.3 vs 13.8
  2. AA-Omniscience · Non-hallucination68.1 vs 49.1

Every shared benchmark32 · grouped by area

Reasoning 3

Coding 6

Agentic 3

GDPVal13.824.3
Terminal-Bench 4.040

Instruction Following 1

IFBench77.167.1

Long Context 1

Math 2

AIME 202698.3 ◆93.6
HMMT Feb. 202694.7 ◆82.6

Multimodal 3

LMArena · Vision1296 ◆◆ 1247
MMMU-Pro82.475.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (google-official, direct), otherwise the lowest tracked offer.

Compare Gemini 3.1 Pro withALL PAIRINGS →