Mistral Large 3 675B Base 2512 is the stronger all-rounder.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Mistral Large 3 675B Base 2512 pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+9points ahead
Where GLM-4.5-Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningMistral Large 3 675B Base 2512
- GPQA DiamondMistral Large 3 675B Base 2512 by 934.943.9+9
Other results20 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- RULER 64KMistral Large 3 675B Base 2512 by 7416.190.1+74
- RULER 128KMistral Large 3 675B Base 2512 by 55.8055.8+55.8
- HumanEvalGLM-4.5-Base by 11.578.266.7+11.5
- MBPP-SanitizedMistral Large 3 675B Base 2512 by 7.476.784.1+7.4
- MMMLUMistral Large 3 675B Base 2512 by 6.279.385.5+6.2
- SimpleQAGLM-4.5-Base by 5.529.323.8+5.5
- MMLU-ProMistral Large 3 675B Base 2512 by 3.763.767.4+3.7
- WinoGrandeGLM-4.5-Base by 3.185.282.1+3.1
- MATHMistral Large 3 675B Base 2512 by 1.96162.9+1.9
- OpenBookQAMistral Large 3 675B Base 2512 by 1.849.651.4+1.8
- MGSMMistral Large 3 675B Base 2512 by 1.681.382.9+1.6
- Global-MMLU-LiteMistral Large 3 675B Base 2512 by 1.585.887.3+1.5
- GPQA (unspecified)Mistral Large 3 675B Base 2512 by 1.433.534.9+1.4
- HellaSwagGLM-4.5-Base by 1.390.288.9+1.3
- GSM8KMistral Large 3 675B Base 2512 by 1.190.191.2+1.1
- RACEMistral Large 3 675B Base 2512 by 1.192.293.3+1.1
- ARC-ChallengeMistral Large 3 675B Base 2512 by 196.397.3+1
- AGIEval-Entie70.169.3tie
- MMLUtie87.787.3tie
- PIQAtie84.784.8tie
Questions people ask
Which is better, GLM-4.5-Base or Mistral Large 3 675B Base 2512?
Mistral Large 3 675B Base 2512 wins the one area where both have results: reasoning. GLM-4.5-Base wins none.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 20 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.