Kimi K2 Base is the stronger all-rounder.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Kimi K2 Base pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4.2points ahead
Where Mistral Large 3 675B Base 2512 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results20 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- RULER 128KKimi K2 Base by 32.888.655.8+32.8
- MBPP-SanitizedMistral Large 3 675B Base 2512 by 1272.184.1+12
- HumanEvalKimi K2 Base by 11.578.266.7+11.5
- SimpleQAKimi K2 Base by 11.535.323.8+11.5
- GPQA (unspecified)Kimi K2 Base by 8.243.134.9+8.2
- MMMLUMistral Large 3 675B Base 2512 by 7.977.685.5+7.9
- MATHKimi K2 Base by 7.370.262.9+7.3
- HellaSwagKimi K2 Base by 5.794.688.9+5.7
- RULER 64KKimi K2 Base by 3.793.890.1+3.7
- AGIEval-EnKimi K2 Base by 3.372.569.3+3.3
- WinoGrandeKimi K2 Base by 3.285.382.1+3.2
- MGSMKimi K2 Base by 2.385.282.9+2.3
- MMLU-ProKimi K2 Base by 1.869.267.4+1.8
- Global-MMLU-LiteMistral Large 3 675B Base 2512 by 1.785.687.3+1.7
- RACEMistral Large 3 675B Base 2512 by 1.39293.3+1.3
- ARC-ChallengeMistral Large 3 675B Base 2512 by 1.196.297.3+1.1
- GSM8Ktie92.191.2tie
- OpenBookQAtie50.851.4tie
- MMLUtie87.887.3tie
- PIQAtie84.984.8tie
Questions people ask
Which is better, Kimi K2 Base or Mistral Large 3 675B Base 2512?
Kimi K2 Base wins the one area where both have results: reasoning. Mistral Large 3 675B Base 2512 wins none.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 20 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.