MiMo V2.5
MiMo V2.5 is capable in long context and multimodal tasks; and behind the leaders in coding, reasoning, instruction following, and agentic tasks. Too few results yet to rate safety, math, multilingual tasks, or factuality.
- Price per million tokens
$0.14input$0.28output
Price from Artificial Analysis · 4 providers tracked · All prices
Cheaper than 82% of 320 priced models · 3:1 input-to-output blend, log scale - Evidence
51results on39benchmarks
- 3 independently verified
- 17 aggregator
- 8 vendor-reported
- 23 cross-referenced
From 7 sources · latest Sep 27, 2026 · How verification works
- API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Capability profile
Bars show the model's median result as a share of the leading model's, per capability.
Too few results to rate: SafetyMathMultilingualFactuality
MiMo V2.5 benchmark results
51 results on 39 benchmarks, grouped by capability. Every score links to its source; a bar is the result as a share of the capability leader's.
Long Context
Capable1 of 3 ranked benchmarks measuredFull long context ranking- AA-LCR73.0082%aa_lcr
Multimodal
Capable3 of 6 ranked benchmarks measuredFull multimodal ranking- 87.7098%
- CharXiv (reasoning)81.0087%CharXiv-R
- MMMU-Pro75.4386%aa_mmmu_pro
4 more multimodal results
- 1246.69—
- 1237—
- 77.90—
- MMMU-Pro75.40—MMMU Pro (Standard 10)
Coding
Limited6 of 10 ranked benchmarks measuredFull coding ranking- SWE-bench Verified71.0074%SWEBench Verified
- Terminal-Bench 2.163.6770%terminalbenchV21
- SciCode43.8766%aa_scicode
- Terminal-Bench Hard41.6763%aa_terminalbench_hard
- 1437.6063%
- SWE-bench Pro56.1062%SWE-Bench Pro
2 more coding results
- SciCode43.10—SciCode
- Terminal-Bench 2.163.70—Terminal Bench 2.1
Reasoning
Limited3 of 6 ranked benchmarks measuredFull reasoning ranking- GPQA Diamond84.9588%gpqa
- Humanity's Last Exam27.2042%aa_hle
- 3.7112%
4 more reasoning results
- CritPt4.00—CritPt
- CritPt3.70—CritPt
- 84.90—
- Humanity's Last Exam25.20—HLE text only
Instruction Following
Limited1 of 3 ranked benchmarks measuredFull instruction following ranking- IFBench67.1481%aa_ifbench
1 more instruction following result
- 67.10—
Agentic
Limited3 of 7 ranked benchmarks measuredFull agentic ranking- 24.3136%
- τ-Bench V3 · Banking8.6617%tauBanking
- Terminal-Bench 4.00.000%
Math
Not enough data0 of 5 ranked benchmarks measuredFull math ranking- HMMT Feb 202682.60—HMMT Feb 2026
- AIME 202693.60—AIME 2026
Factuality
Not enough data0 of 4 ranked benchmarks measuredFull factuality ranking- AA-Omniscience · Accuracy16.75—omniscienceAccuracy
- AA-Omniscience · Non-hallucination68.07—omniscienceNonHallucination
- SimpleQA Verified16.10—SimpleQA Verified
More results
- AA-Omniscience-9.83—aa_omniscience
- τ²-Bench Telecom (AA run)90.64—aa_tau2
- AA Intelligence25.17—aa_intelligence_index
- 36.73—
- 87.20—
- 65.80—
12 more results
- Claw Eval (pass@3)63.20—Claw-Eval
- 1265.10—
- 1145.00—
- 83.50—
- HLE (with tools)40.00—HLE with tools
- 73.60—
- StrongREJECT99.30—StrongREJECT
- SWEBench Pro Public56.10—SWEBench Pro public
- 49.10—
- 86.40—
- τ³-Bench Banking9.00—τ³-Banking
- τ³-Bench Banking6.60—Tau 3 Banking
About this record
Verification: 51 scores · 3 independently verified · 17 aggregator-attributed · 23 vendor cross-reference · 8 vendor-reported. How these tiers are assigned
- Tracked since
- Sep 23, 2026
- Newest source mention
- Sep 23, 2026