Kimi K2 Base is the stronger all-rounder.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Kimi K2 Base pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+12.8points ahead
Where DeepSeek-V2 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HumanEvalKimi K2 Base by 34.943.378.2+34.9
- MATHKimi K2 Base by 26.843.470.2+26.8
- Chinese SimpleQA (C-SimpleQA)Kimi K2 Base by 26.148.574.6+26.1
- EvalPlusKimi K2 Base by 25.35580.3+25.3
- MBPPKimi K2 Base by 246589+24
- MGSMKimi K2 Base by 21.663.685.2+21.6
- CRUXEval-O (output prediction)Kimi K2 Base by 19.849.869.6+19.8
- MMLU-ProKimi K2 Base by 17.851.469.2+17.8
- CRUXEval-I (input prediction)Kimi K2 Base by 15.552.568+15.5
- MMMLUKimi K2 Base by 13.66477.6+13.6
- GSM8KKimi K2 Base by 12.979.292.1+12.9
- MMLU-ReduxKimi K2 Base by 12.377.990.2+12.3
- MMLUKimi K2 Base by 12.275.687.8+12.2
- CMathKimi K2 Base by 12.178.790.8+12.1
- C-EvalKimi K2 Base by 11.181.492.5+11.1
- BBHKimi K2 Base by 9.978.888.7+9.9
- HellaSwagKimi K2 Base by 7.587.194.6+7.5
- CMMLUKimi K2 Base by 6.98490.9+6.9
- TriviaQAKimi K2 Base by 5.279.985.1+5.2
- ARC-ChallengeKimi K2 Base by 492.296.2+4
- DROPKimi K2 Base by 3.580.183.6+3.5
- PIQAKimi K2 Base by 183.984.9+1
- WinoGrandeDeepSeek-V2 by 186.385.3+1
Questions people ask
Which is better, DeepSeek-V2 or Kimi K2 Base?
Kimi K2 Base wins the one area where both have results: reasoning. DeepSeek-V2 wins none.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.