Across 6 shared benchmarks, Llama 3.1 Instruct 405B scores higher on 4 and Llama 3.1 Nemotron Nano 8B V1 on 2. The widest gap is math_500_em, where Llama 3.1 Nemotron Nano 8B V1 scores 95.4 against 73.8.
| Benchmark | Llama 3.1 Instruct 405B | Llama 3.1 Nemotron Nano 8B V1 |
|---|---|---|
| BFCL v2 | 81.1 | 63.6 |
| GPQA Diamond | 51.5 | 54.1 |
| ifeval | 88.6 | 79.3 |
| math_500_em | 73.8 | 95.4 |
| MBPP | 73.4 | 0.8 |
| mt_bench | 9.1 | 0.8 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.