Across 15 shared benchmarks, ERNIE 4.5 scores higher on 1 and Kimi K2 Base on 14. The widest gap is MMLU-Pro, where Kimi K2 Base scores 69.2 against 16.
| Benchmark | ERNIE 4.5 | Kimi K2 Base |
|---|---|---|
| arc_challenge | 40.6 | 96.7 |
| bbh | 30.4 | 88.7 |
| C-Eval | 40.7 | 92.5 |
| cmmlu | 39.8 | 90.9 |
| DROP | 28.6 | 86.4 |
| GPQA Diamond | 74 | 48.1 |
| GSM8K | 25.2 | 93.5 |
| hellaswag | 33 | 94.6 |
| HumanEval+ | 25 | 84.8 |
| MBPP+ | 40.2 | 73.8 |
| mmlu_redux | 43.2 | 90.2 |
| MMLU-Pro | 16 | 69.2 |
| PIQA | 55.2 | 85.5 |
| simpleqa | 1.8 | 35.3 |
| winogrande | 51.3 | 85.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.