Llama 3.1 Instruct 405B vs Qwen2.5 Instruct 72B
Wins 4 of 5 areas
Coding · Reasoning · Facts · Following instructions
Wins 0 of 5 areas
—
Llama 3.1 Instruct 405B is the stronger all-rounder.
Scores updated · 89 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20Llama 3.1 Instruct 405B2 of 3 tests · 1 tie
- FactsGetting facts right instead of making them up20Llama 3.1 Instruct 405B2 of 2 tests
- ReasoningHard problems that need careful thinking10Llama 3.1 Instruct 405B1 of 3 tests · 2 ties
- Following instructionsDoing exactly what it is asked10Llama 3.1 Instruct 405B1 of 1 test
- Long documentsFinding answers in very long texts11Even1 each
Following instructions rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Llama 3.1 Instruct 405B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+33.1points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+5.7points ahead
- Reasons across sets of long documentsAA-LCR+5.3points ahead
Where Qwen2.5 Instruct 72B pulls ahead
- Questions about very long textsLongBench v2+3.3points ahead
Every test, side by side
All 89 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingLlama 3.1 Instruct 405B
- SciCodeLlama 3.1 Instruct 405B by 3.229.926.7+3.2
- Terminal-Bench HardLlama 3.1 Instruct 405B by 2.36.84.5+2.3
- SWE-bench Verifiedtie24.523.8tie
ReasoningLlama 3.1 Instruct 405B
- GPQA DiamondLlama 3.1 Instruct 405B by 2.451.549.1+2.4
- Humanity's Last Examtie43.6tie
- CritPttie00tie
FactsLlama 3.1 Instruct 405B
- AA-Omniscience · Non-hallucinationLlama 3.1 Instruct 405B by 33.147.614.5+33.1
- AA-Omniscience · AccuracyLlama 3.1 Instruct 405B by 5.723.217.5+5.7
Long documentsEven
- AA-LCRLlama 3.1 Instruct 405B by 5.325.320+5.3
- LongBench v2Qwen2.5 Instruct 72B by 3.336.139.4+3.3
Following instructionsLlama 3.1 Instruct 405B
- IFBenchLlama 3.1 Instruct 405B by 2.13936.9+2.1
Other results78 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AlignBenchQwen2.5 Instruct 72B by 75.6681.6+75.6
- AA-OmniscienceLlama 3.1 Instruct 405B by 36-17.1-53.1+36
- LiveCodeBenchQwen2.5 Instruct 72B by 27.827.755.5+27.8
- MBPPQwen2.5 Instruct 72B by 19.868.488.2+19.8
- C-EvalQwen2.5 Instruct 72B by 16.772.589.2+16.7
- CMMLUQwen2.5 Instruct 72B by 15.873.789.5+15.8
- CMMLU (Acc.)Qwen2.5 Instruct 72B by 15.873.789.5+15.8
- τ²-Bench Telecom (AA run)Qwen2.5 Instruct 72B by 15.51934.5+15.5
- MGSMLlama 3.1 Instruct 405B by 15.491.676.2+15.4
- AGIEvalQwen2.5 Instruct 72B by 15.260.675.8+15.2
- MATHQwen2.5 Instruct 72B by 14.673.888.4+14.6
- DROP (3-shot F1)Llama 3.1 Instruct 405B by 1288.776.7+12
- Arena HardQwen2.5 Instruct 72B by 11.969.381.2+11.9
- MMLULlama 3.1 Instruct 405B by 11.688.677+11.6
- TriviaQA (EM)Llama 3.1 Instruct 405B by 10.882.771.9+10.8
- HarmBenchQwen2.5 Instruct 72B by 10.162.772.8+10.1
- CCPMQwen2.5 Instruct 72B by 9.978.688.5+9.9
- naturalquestions_closedbookLlama 3.1 Instruct 405B by 9.745.635.9+9.7
- NaturalQuestionsLlama 3.1 Instruct 405B by 8.341.533.2+8.3
- NaturalQuestions (EM)Llama 3.1 Instruct 405B by 8.341.533.2+8.3
- DROPLlama 3.1 Instruct 405B by 8.184.876.7+8.1
- SimpleQALlama 3.1 Instruct 405B by 817.19.1+8
- CMath (EM)Qwen2.5 Instruct 72B by 7.277.384.5+7.2
- RACE-HighLlama 3.1 Instruct 405B by 6.556.850.3+6.5
- MGSM (EM)Qwen2.5 Instruct 72B by 6.369.976.2+6.3
- MATH-500 (EM)Qwen2.5 Instruct 72B by 6.273.880+6.2
- RACE-MiddleLlama 3.1 Instruct 405B by 6.174.268.1+6.1
- MMLU-Pro (Acc.)Qwen2.5 Instruct 72B by 5.552.858.3+5.5
- DROP (F1)Llama 3.1 Instruct 405B by 5.48680.6+5.4
- MATH (EM)Qwen2.5 Instruct 72B by 5.44954.4+5.4
- GSM8K (EM)Qwen2.5 Instruct 72B by 4.883.588.3+4.8
- IFEvalLlama 3.1 Instruct 405B by 4.588.684.1+4.5
- HellaSwagLlama 3.1 Instruct 405B by 4.489.284.8+4.4
- HellaSwag (Acc.)Llama 3.1 Instruct 405B by 4.489.284.8+4.4
- MBPP+ (EvalPlus-augmented)Qwen2.5 Instruct 72B by 47377+4
- PIQALlama 3.1 Instruct 405B by 3.385.982.6+3.3
- PIQA (Acc.)Llama 3.1 Instruct 405B by 3.385.982.6+3.3
- anthropic_red_teamQwen2.5 Instruct 72B by 3.196.599.6+3.1
- BBHLlama 3.1 Instruct 405B by 3.182.979.8+3.1
- BBH (EM)Llama 3.1 Instruct 405B by 3.182.979.8+3.1
- C3Llama 3.1 Instruct 405B by 379.776.7+3
- C3 (Acc.)Llama 3.1 Instruct 405B by 379.776.7+3
- WinoGrandeLlama 3.1 Instruct 405B by 2.985.282.3+2.9
- WinoGrande (Acc.)Llama 3.1 Instruct 405B by 2.985.282.3+2.9
- Artificial Analysis Coding IndexLlama 3.1 Instruct 405B by 2.614.511.9+2.6
- LiveCodeBench-Base (Pass@1)Llama 3.1 Instruct 405B by 2.615.512.9+2.6
- ARC-ChallengeLlama 3.1 Instruct 405B by 2.496.994.5+2.4
- HumanEvalLlama 3.1 Instruct 405B by 2.48986.6+2.4
- MMLU-ProLlama 3.1 Instruct 405B by 2.273.371.1+2.2
- OpenBookQAQwen2.5 Instruct 72B by 2.29496.2+2.2
- Chinese SimpleQA (C-SimpleQA)Llama 3.1 Instruct 405B by 250.448.4+2
- XSTestQwen2.5 Instruct 72B by 295.997.9+2
- MMLU-Redux (Acc.)Qwen2.5 Instruct 72B by 1.981.383.2+1.9
- Aider-Edit (Acc.)Qwen2.5 Instruct 72B by 1.563.965.4+1.5
- simple_safety_testsQwen2.5 Instruct 72B by 1.298.8100+1.2
- GSM8KLlama 3.1 Instruct 405B by 196.895.8+1
- MMMLUQwen2.5 Instruct 72B by 173.874.8+1
- MMMLU-non-English (Acc.)Qwen2.5 Instruct 72B by 173.874.8+1
- bbqtie94.595.4tie
- MT-Benchtie8.59.3tie
- ARC-Challenge (Acc.)tie95.394.5tie
- IFEval (avg)tie86.487.2tie
- CRUXEval-I (input prediction)tie58.559.1tie
- MMLU (Acc.)tie84.485tie
- MMLU-Reduxtie86.286.8tie
- CLUEWSCtie8382.5tie
- Codeforcestie25.324.8tie
- AA Intelligencetie7.37.7tie
- AIR-Bench 2024tie58.659tie
- NarrativeQAtie74.974.5tie
- CMRCtie7675.8tie
- FRAMES (Acc.)tie7069.8tie
- HumanEval-Mul (Pass@1)tie77.277.3tie
- Pile-test (BPB)lower is bettertie0.50.6tie
- The Pile (Test, BPB)lower is bettertie0.50.6tie
- AIME 2024tie23.323.3tie
- ARC-Easy (Acc.)tie98.498.4tie
- CRUXEval-O (output prediction)tie59.959.9tie
Questions people ask
Which is better, Llama 3.1 Instruct 405B or Qwen2.5 Instruct 72B?
Llama 3.1 Instruct 405B wins four of the five areas where both have results: coding, reasoning, facts and following instructions. Qwen2.5 Instruct 72B wins none. They are level on long documents.
Which is better for coding?
Llama 3.1 Instruct 405B. It wins 2 of the 3 coding tests both models report; Qwen2.5 Instruct 72B wins none, and 1 is a tie.
How do you compare the two?
We use the 89 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 78 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.