Llama 3.1 Instruct 405B vs Llama 3 Instruct 8B
Wins 4 of 5 areas
Coding · Facts · Long documents · Following instructions
Wins 0 of 5 areas
—
Llama 3.1 Instruct 405B is the stronger all-rounder.
Scores updated · 36 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20Llama 3.1 Instruct 405B2 of 2 tests
- FactsGetting facts right instead of making them up20Llama 3.1 Instruct 405B2 of 2 tests
- Long documentsFinding answers in very long texts10Llama 3.1 Instruct 405B1 of 1 test
- Following instructionsDoing exactly what it is asked10Llama 3.1 Instruct 405B1 of 1 test
- ReasoningHard problems that need careful thinking11Even1 each · 1 tie
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Llama 3.1 Instruct 405B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+37.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+21.9points ahead
- Code for real scientific research problemsSciCode+18points ahead
Where Llama 3 Instruct 8B pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.1points ahead
Every test, side by side
All 36 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingLlama 3.1 Instruct 405B
- SciCodeLlama 3.1 Instruct 405B by 1829.911.9+18
- Terminal-Bench HardLlama 3.1 Instruct 405B by 6.86.80+6.8
ReasoningEven
- GPQA DiamondLlama 3.1 Instruct 405B by 21.951.529.6+21.9
- Humanity's Last ExamLlama 3 Instruct 8B by 1.145.1+1.1
- CritPttie00tie
FactsLlama 3.1 Instruct 405B
- AA-Omniscience · Non-hallucinationLlama 3.1 Instruct 405B by 37.447.610.2+37.4
- AA-Omniscience · AccuracyLlama 3.1 Instruct 405B by 12.823.210.4+12.8
Long documentsLlama 3.1 Instruct 405B
- AA-LCRLlama 3.1 Instruct 405B by 25.325.30+25.3
Following instructionsLlama 3.1 Instruct 405B
- IFBenchLlama 3.1 Instruct 405B by 14.43924.6+14.4
Other results27 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceLlama 3.1 Instruct 405B by 53-17.1-70.1+53
- GSM8KLlama 3.1 Instruct 405B by 46.996.849.9+46.9
- MMLU-ProLlama 3.1 Instruct 405B by 37.973.335.4+37.9
- MATHLlama 3.1 Instruct 405B by 34.773.839.1+34.7
- HumanEvalLlama 3.1 Instruct 405B by 28.68960.4+28.6
- MMLULlama 3.1 Instruct 405B by 28.488.660.1+28.4
- BBHLlama 3.1 Instruct 405B by 25.282.957.7+25.2
- C-EvalLlama 3.1 Instruct 405B by 2372.549.5+23
- CMMLULlama 3.1 Instruct 405B by 22.973.750.8+22.9
- WinoGrandeLlama 3.1 Instruct 405B by 20.285.265+20.2
- τ²-Bench Telecom (AA run)Llama 3.1 Instruct 405B by 19190+19
- HellaSwagLlama 3.1 Instruct 405B by 18.189.271.1+18.1
- bbqLlama 3.1 Instruct 405B by 1894.576.5+18
- OpenBookQALlama 3.1 Instruct 405B by 17.49476.6+17.4
- ARC-ChallengeLlama 3.1 Instruct 405B by 14.196.982.8+14.1
- AIR-Bench 2024Llama 3 Instruct 8B by 12.358.670.9+12.3
- Artificial Analysis Coding IndexLlama 3.1 Instruct 405B by 10.514.54+10.5
- PIQALlama 3.1 Instruct 405B by 10.285.975.7+10.2
- HarmBenchLlama 3 Instruct 8B by 1062.772.7+10
- naturalquestions_closedbookLlama 3.1 Instruct 405B by 7.845.637.8+7.8
- AA Agentic IndexLlama 3.1 Instruct 405B by 6.36.30+6.3
- AA IntelligenceLlama 3.1 Instruct 405B by 2.57.34.8+2.5
- anthropic_red_teamLlama 3 Instruct 8B by 2.396.598.8+2.3
- MBPPtie68.467.7tie
- NarrativeQAtie74.975.4tie
- simple_safety_teststie98.899.3tie
- XSTesttie95.995.6tie
Questions people ask
Which is better, Llama 3.1 Instruct 405B or Llama 3 Instruct 8B?
Llama 3.1 Instruct 405B wins four of the five areas where both have results: coding, facts, long documents and following instructions. Llama 3 Instruct 8B wins none. They are level on reasoning.
Which is better for coding?
Llama 3.1 Instruct 405B. It wins 2 of the 2 coding tests both models report; Llama 3 Instruct 8B wins none.
How do you compare the two?
We use the 36 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 27 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.