MiMo V2.5 Base is the stronger all-rounder.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where MiMo V2.5 Base pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+10points ahead
- Recent programming contest problemsLiveCodeBench v6+9.2points ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+2.6points ahead
Where Kimi K2 Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingMiMo V2.5 Base
- LiveCodeBench v6MiMo V2.5 Base by 9.226.335.5+9.2
- SWE-bench VerifiedMiMo V2.5 Base by 2.628.230.8+2.6
Other results18 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HumanEval+Kimi K2 Base by 13.584.871.3+13.5
- GSM8KKimi K2 Base by 8.892.183.3+8.8
- HellaSwagKimi K2 Base by 694.688.6+6
- aime_2024_2025_combinedMiMo V2.5 Base by 5.331.636.9+5.3
- TriviaQAKimi K2 Base by 4.485.180.7+4.4
- C-EvalKimi K2 Base by 3.992.588.6+3.9
- MMLU-ProKimi K2 Base by 3.469.265.8+3.4
- GlobalMMLUKimi K2 Base by 3.380.777.4+3.3
- MBPP+Kimi K2 Base by 2.973.870.9+2.9
- MBPP+ (EvalPlus-augmented)Kimi K2 Base by 2.973.870.9+2.9
- CMMLUKimi K2 Base by 2.790.988.2+2.7
- MATHKimi K2 Base by 2.570.267.7+2.5
- BBHKimi K2 Base by 1.588.787.2+1.5
- MMLUKimi K2 Base by 1.587.886.3+1.5
- WinoGrandetie85.384.7tie
- MMLU-Reduxtie90.289.8tie
- ARC-Challengetie96.296.5tie
- DROPtie83.683.7tie
Questions people ask
Which is better, Kimi K2 Base or MiMo V2.5 Base?
MiMo V2.5 Base wins all two areas where both have results: coding and reasoning. Kimi K2 Base wins none.
Which is better for coding?
MiMo V2.5 Base. It wins 2 of the 2 coding tests both models report; Kimi K2 Base wins none.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 18 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.