Kimi K2 Base is the stronger all-rounder.
Scores updated · 13 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Kimi K2 Base pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+10.2points ahead
Where Qwen2 72B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 13 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results12 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ARC-ChallengeKimi K2 Base by 27.396.268.9+27.3
- MATHKimi K2 Base by 19.170.251.1+19.1
- EvalPlusKimi K2 Base by 14.980.365.4+14.9
- HumanEvalKimi K2 Base by 13.678.264.6+13.6
- MMLU-ProKimi K2 Base by 13.669.255.6+13.6
- MBPPKimi K2 Base by 12.18976.9+12.1
- HellaSwagKimi K2 Base by 794.687.6+7
- BBHKimi K2 Base by 6.388.782.4+6.3
- MMLUKimi K2 Base by 3.687.884.2+3.6
- GSM8KKimi K2 Base by 2.692.189.5+2.6
- C-EvalKimi K2 Base by 1.592.591+1.5
- CMMLUtie90.990.1tie
Questions people ask
Which is better, Kimi K2 Base or Qwen2 72B?
Kimi K2 Base wins the one area where both have results: reasoning. Qwen2 72B wins none.
How do you compare the two?
We use the 13 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 12 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.