Across 14 shared benchmarks, ERNIE 4.5 scores higher on 1 and Llama 3.1 Instruct 405B on 13. The widest gap is MMLU-Pro, where Llama 3.1 Instruct 405B scores 73.4 against 16.
| Benchmark | ERNIE 4.5 | Llama 3.1 Instruct 405B |
|---|---|---|
| arc_challenge | 40.6 | 96.9 |
| bbh | 30.4 | 85.9 |
| C-Eval | 40.7 | 72.5 |
| CLUEWSC | 48.6 | 84.7 |
| cmmlu | 39.8 | 73.7 |
| DROP | 28.6 | 84.8 |
| GPQA Diamond | 74 | 51.5 |
| GSM8K | 25.2 | 96.8 |
| hellaswag | 33 | 89.2 |
| mmlu_redux | 43.2 | 86.2 |
| MMLU-Pro | 16 | 73.4 |
| PIQA | 55.2 | 85.9 |
| simpleqa | 1.8 | 23.2 |
| winogrande | 51.3 | 86.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.