MiMo V2.5 Pro Base is the stronger all-rounder.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where MiMo V2.5 Pro Base pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+18.6points ahead
- Recent programming contest problemsLiveCodeBench v6+13.3points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+7.5points ahead
Where Kimi K2 Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingMiMo V2.5 Pro Base
- LiveCodeBench v6MiMo V2.5 Pro Base by 13.326.339.6+13.3
- SWE-bench VerifiedMiMo V2.5 Pro Base by 7.528.235.7+7.5
ReasoningMiMo V2.5 Pro Base
- GPQA DiamondMiMo V2.5 Pro Base by 18.648.166.7+18.6
Other results18 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MATHMiMo V2.5 Pro Base by 1670.286.2+16
- HumanEval+Kimi K2 Base by 9.284.875.6+9.2
- GSM8KMiMo V2.5 Pro Base by 7.592.199.6+7.5
- aime_2024_2025_combinedMiMo V2.5 Pro Base by 5.731.637.3+5.7
- HellaSwagKimi K2 Base by 4.894.689.8+4.8
- TriviaQAKimi K2 Base by 3.885.181.3+3.8
- GlobalMMLUMiMo V2.5 Pro Base by 2.980.783.6+2.9
- DROPMiMo V2.5 Pro Base by 2.783.686.3+2.7
- MMLU-ReduxMiMo V2.5 Pro Base by 2.690.292.8+2.6
- MMLUMiMo V2.5 Pro Base by 1.687.889.4+1.6
- ARC-ChallengeMiMo V2.5 Pro Base by 196.297.2+1
- C-EvalKimi K2 Base by 192.591.5+1
- CMMLUtie90.990.2tie
- MMLU-Protie69.268.5tie
- BBHtie88.788.4tie
- MBPP+tie73.874.1tie
- MBPP+ (EvalPlus-augmented)tie73.874.1tie
- WinoGrandetie85.385.6tie
Questions people ask
Which is better, Kimi K2 Base or MiMo V2.5 Pro Base?
MiMo V2.5 Pro Base wins all two areas where both have results: coding and reasoning. Kimi K2 Base wins none.
Which is better for coding?
MiMo V2.5 Pro Base. It wins 2 of the 2 coding tests both models report; Kimi K2 Base wins none.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 18 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.