Llama 3.1 Instruct 405B vs Qwen 2.5 Coder 32B Instruct
Wins 2 of 2 areas
Coding · Reasoning
Wins 0 of 2 areas
—
Llama 3.1 Instruct 405B is the stronger all-rounder.
Scores updated · 14 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Llama 3.1 Instruct 405B pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+15.5points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+9.8points ahead
- Code for real scientific research problemsSciCode+2.8points ahead
Where Qwen 2.5 Coder 32B Instruct pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 14 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingLlama 3.1 Instruct 405B
- SWE-bench VerifiedLlama 3.1 Instruct 405B by 15.524.59+15.5
- SciCodeLlama 3.1 Instruct 405B by 2.829.927.1+2.8
ReasoningLlama 3.1 Instruct 405B
- GPQA DiamondLlama 3.1 Instruct 405B by 9.851.541.7+9.8
- Humanity's Last Examtie43.5tie
Other results10 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ARC-ChallengeLlama 3.1 Instruct 405B by 26.496.970.5+26.4
- MMLU-ProLlama 3.1 Instruct 405B by 22.973.350.4+22.9
- MBPPQwen 2.5 Coder 32B Instruct by 21.868.490.2+21.8
- MMLU-ReduxLlama 3.1 Instruct 405B by 8.786.277.5+8.7
- HellaSwagLlama 3.1 Instruct 405B by 6.289.283+6.2
- GSM8KLlama 3.1 Instruct 405B by 5.796.891.1+5.7
- WinoGrandeLlama 3.1 Instruct 405B by 4.485.280.8+4.4
- HumanEvalQwen 2.5 Coder 32B Instruct by 3.78992.7+3.7
- LiveCodeBenchQwen 2.5 Coder 32B Instruct by 3.727.731.4+3.7
- AA Intelligencetie7.36.7tie
Questions people ask
Which is better, Llama 3.1 Instruct 405B or Qwen 2.5 Coder 32B Instruct?
Llama 3.1 Instruct 405B wins all two areas where both have results: coding and reasoning. Qwen 2.5 Coder 32B Instruct wins none.
Which is better for coding?
Llama 3.1 Instruct 405B. It wins 2 of the 2 coding tests both models report; Qwen 2.5 Coder 32B Instruct wins none.
How do you compare the two?
We use the 14 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 10 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.