Llama 3.1 Instruct 405B is the stronger all-rounder.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Llama 3.1 Instruct 405B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+16.6points ahead
Where GLM-4.5-Base pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningLlama 3.1 Instruct 405B
- GPQA DiamondLlama 3.1 Instruct 405B by 16.634.951.5+16.6
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- OpenBookQALlama 3.1 Instruct 405B by 44.449.694+44.4
- Chinese SimpleQA (C-SimpleQA)GLM-4.5-Base by 18.168.550.4+18.1
- C-EvalGLM-4.5-Base by 13.385.872.5+13.3
- MBPPGLM-4.5-Base by 13.281.668.4+13.2
- CMMLUGLM-4.5-Base by 12.886.573.7+12.8
- MATHLlama 3.1 Instruct 405B by 12.86173.8+12.8
- SimpleQAGLM-4.5-Base by 12.229.317.1+12.2
- HumanEvalLlama 3.1 Instruct 405B by 10.878.289+10.8
- MGSMLlama 3.1 Instruct 405B by 10.381.391.6+10.3
- CRUXEval-I (input prediction)GLM-4.5-Base by 1068.558.5+10
- MMLU-ProLlama 3.1 Instruct 405B by 9.663.773.3+9.6
- CRUXEval-O (output prediction)GLM-4.5-Base by 7.967.859.9+7.9
- GSM8KLlama 3.1 Instruct 405B by 6.790.196.8+6.7
- MMMLUGLM-4.5-Base by 5.579.373.8+5.5
- MBPP+ (EvalPlus-augmented)GLM-4.5-Base by 5.17873+5.1
- BBHGLM-4.5-Base by 3.386.282.9+3.3
- DROPLlama 3.1 Instruct 405B by 1.982.984.8+1.9
- PIQALlama 3.1 Instruct 405B by 1.284.785.9+1.2
- HellaSwagGLM-4.5-Base by 190.289.2+1
- MMLUtie87.788.6tie
- ARC-Challengetie96.396.9tie
- MMLU-Reduxtie86.686.2tie
- WinoGrandetie85.285.2tie
Questions people ask
Which is better, GLM-4.5-Base or Llama 3.1 Instruct 405B?
Llama 3.1 Instruct 405B wins the one area where both have results: reasoning. GLM-4.5-Base wins none.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.