Across 14 shared benchmarks, ERNIE 4.5 scores higher on 1 and Qwen2.5 Instruct 72B on 13. The widest gap is simpleqa, where Qwen2.5 Instruct 72B scores 10.3 against 1.8.
| Benchmark | ERNIE 4.5 | Qwen2.5 Instruct 72B |
|---|---|---|
| arc_challenge | 40.6 | 94.5 |
| bbh | 30.4 | 79.8 |
| C-Eval | 40.7 | 89.2 |
| CLUEWSC | 48.6 | 91.4 |
| cmmlu | 39.8 | 89.5 |
| DROP | 28.6 | 76.7 |
| GPQA Diamond | 74 | 49.1 |
| GSM8K | 25.2 | 95.8 |
| hellaswag | 33 | 84.8 |
| mmlu_redux | 43.2 | 86.8 |
| MMLU-Pro | 16 | 71.6 |
| PIQA | 55.2 | 82.6 |
| simpleqa | 1.8 | 10.3 |
| winogrande | 51.3 | 82.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.