Claude 3.5 Sonnet vs Llama 3.1 Instruct 405B
Wins 3 of 3 areas
Coding · Reasoning · Long documents
Wins 0 of 3 areas
—
Claude 3.5 Sonnet is the stronger all-rounder.
Scores updated · 45 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude 3.5 Sonnet pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+24.5points ahead
- Questions about very long textsLongBench v2+4.9points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4.5points ahead
Where Llama 3.1 Instruct 405B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 45 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude 3.5 Sonnet
- SWE-bench VerifiedClaude 3.5 Sonnet by 24.54924.5+24.5
- SciCodeClaude 3.5 Sonnet by 1.731.629.9+1.7
ReasoningClaude 3.5 Sonnet
- GPQA DiamondClaude 3.5 Sonnet by 4.55651.5+4.5
- SimpleBenchClaude 3.5 Sonnet by 4.527.523+4.5
- Humanity's Last Examtie3.24tie
Long documentsClaude 3.5 Sonnet
- LongBench v2Claude 3.5 Sonnet by 4.94136.1+4.9
Other results39 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- CodeforcesClaude 3.5 Sonnet by 691.7 rating points71725.3+691.7 rating
- HarmBenchClaude 3.5 Sonnet by 35.498.162.7+35.4
- AIR-Bench 2024Claude 3.5 Sonnet by 27.385.958.6+27.3
- Aider-Edit (Acc.)Claude 3.5 Sonnet by 20.384.263.9+20.3
- AlpacaEval 2 LCClaude 3.5 Sonnet by 13.152.439.3+13.1
- Artificial Analysis Coding IndexClaude 3.5 Sonnet by 11.52614.5+11.5
- SimpleQAClaude 3.5 Sonnet by 11.328.417.1+11.3
- BBHClaude 3.5 Sonnet by 10.293.182.9+10.2
- Arena HardClaude 3.5 Sonnet by 9.979.269.3+9.9
- AIME 2024Llama 3.1 Instruct 405B by 7.31623.3+7.3
- LiveCodeBenchClaude 3.5 Sonnet by 5.132.827.7+5.1
- Chinese SimpleQA (C-SimpleQA)Claude 3.5 Sonnet by 555.450.4+5
- HumanEvalClaude 3.5 Sonnet by 4.793.789+4.7
- naturalquestions_closedbookClaude 3.5 Sonnet by 4.650.245.6+4.6
- HumanEval-Mul (Pass@1)Claude 3.5 Sonnet by 4.581.777.2+4.5
- MATH-500 (EM)Claude 3.5 Sonnet by 4.578.373.8+4.5
- MMLU-ProClaude 3.5 Sonnet by 4.377.673.3+4.3
- C-EvalClaude 3.5 Sonnet by 4.276.772.5+4.2
- IFEval (avg)Claude 3.5 Sonnet by 3.790.186.4+3.7
- anthropic_red_teamClaude 3.5 Sonnet by 3.399.896.5+3.3
- OpenBookQAClaude 3.5 Sonnet by 3.297.294+3.2
- DROP (F1)Claude 3.5 Sonnet by 2.888.886+2.8
- MATHLlama 3.1 Instruct 405B by 2.771.173.8+2.7
- MMLU-ReduxClaude 3.5 Sonnet by 2.788.986.2+2.7
- FRAMES (Acc.)Claude 3.5 Sonnet by 2.572.570+2.5
- CLUEWSCClaude 3.5 Sonnet by 2.485.483+2.4
- DROPClaude 3.5 Sonnet by 2.387.184.8+2.3
- IFEvalLlama 3.1 Instruct 405B by 2.186.588.6+2.1
- MBPP+ (EvalPlus-augmented)Claude 3.5 Sonnet by 2.175.173+2.1
- simple_safety_testsClaude 3.5 Sonnet by 1.210098.8+1.2
- bbqtie94.994.5tie
- DROP (3-shot F1)tie88.388.7tie
- GSM8Ktie96.496.8tie
- MT-Benchtie8.88.5tie
- NarrativeQAtie74.674.9tie
- XSTesttie95.695.9tie
- AA Intelligencetie7.27.3tie
- MMLUtie88.788.6tie
- MGSMtie91.691.6tie
Questions people ask
Which is better, Claude 3.5 Sonnet or Llama 3.1 Instruct 405B?
Claude 3.5 Sonnet wins all three areas where both have results: coding, reasoning and long documents. Llama 3.1 Instruct 405B wins none.
Which is better for coding?
Claude 3.5 Sonnet. It wins 2 of the 2 coding tests both models report; Llama 3.1 Instruct 405B wins none.
How do you compare the two?
We use the 45 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 39 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.