Llama 3 (70B) is the stronger all-rounder.
Scores updated · 20 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Llama 3 (70B) pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+6.7points ahead
Where Llama 3 Instruct 8B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 20 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results19 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Multi-MathematicsLlama 3 (70B) by 30.867.136.3+30.8
- GSM8KLlama 3 (70B) by 30.680.549.9+30.6
- MATHLlama 3 (70B) by 27.266.339.1+27.2
- MultiPL-ELlama 3 (70B) by 23.746.322.6+23.7
- BBHLlama 3 (70B) by 23.38157.7+23.3
- TruthfulQALlama 3 Instruct 8B by 17.645.663.2+17.6
- MMLU-ProLlama 3 (70B) by 17.452.835.4+17.4
- HellaSwagLlama 3 (70B) by 16.98871.1+16.9
- OpenBookQALlama 3 (70B) by 16.893.476.6+16.8
- CMMLULlama 3 (70B) by 16.467.250.8+16.4
- C-EvalLlama 3 (70B) by 15.765.249.5+15.7
- EvalPlusLlama 3 (70B) by 14.554.840.3+14.5
- ARC-ChallengeLlama 3 Instruct 8B by 1468.882.8+14
- HumanEvalLlama 3 Instruct 8B by 12.248.260.4+12.2
- Theorem QALlama 3 (70B) by 10.232.322.1+10.2
- naturalquestions_closedbookLlama 3 (70B) by 9.747.537.8+9.7
- MMLULlama 3 (70B) by 9.369.560.1+9.3
- NarrativeQALlama 3 (70B) by 4.479.875.4+4.4
- MBPPLlama 3 (70B) by 2.770.467.7+2.7
Questions people ask
Which is better, Llama 3 (70B) or Llama 3 Instruct 8B?
Llama 3 (70B) wins the one area where both have results: reasoning. Llama 3 Instruct 8B wins none.
How do you compare the two?
We use the 20 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 19 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.