Across 9 shared benchmarks, Gemma 3 1B scores higher on 0 and Llama 3.1 Instruct 405B on 9. The widest gap is MMLU-Pro, where Llama 3.1 Instruct 405B scores 73.4 against 14.7.
| Benchmark | Gemma 3 1B | Llama 3.1 Instruct 405B |
|---|---|---|
| bbh | 39.1 | 85.9 |
| GPQA Diamond | 19.2 | 51.5 |
| GSM8K | 62.8 | 96.8 |
| humaneval | 41.5 | 89 |
| ifeval | 80.2 | 88.6 |
| LiveCodeBench | 1.9 | 30.1 |
| MBPP | 0.4 | 73.4 |
| MMLU-Pro | 14.7 | 73.4 |
| simpleqa | 2.2 | 23.2 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.