Across 20 shared benchmarks, Llama 3 Instruct 8B scores higher on 16 and Mistral 7B on 4. The widest gap is MATH, where Llama 3 Instruct 8B scores 39.1 against 13.1.
| Benchmark | Llama 3 Instruct 8B | Mistral 7B |
|---|---|---|
| arc_challenge | 59.3 | 60 |
| bbh | 57.7 | 56.1 |
| C-Eval | 49.5 | 47.4 |
| EvalPlus | 40.3 | 36.4 |
| GPQA Diamond | 29.6 | 24.7 |
| GSM8K | 56 | 52.2 |
| hellaswag | 82.1 | 83.2 |
| humaneval | 33.5 | 29.3 |
| MATH | 39.1 | 13.1 |
| MBPP | 53.9 | 51.1 |
| mmlu | 66.6 | 64.2 |
| MMLU-Pro | 35.4 | 30.9 |
| Multi-Exam | 52.3 | 47.1 |
| Multi-Mathematics | 36.3 | 26.3 |
| Multi-Translation | 31.9 | 23.3 |
| Multi-Understanding | 68.6 | 63.3 |
| MultiPL-E | 22.6 | 29.4 |
| Theorem QA | 22.1 | 19.2 |
| TruthfulQA | 44 | 42.2 |
| winogrande | 77.4 | 78.4 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.