Across 15 shared benchmarks, DeepSeek-V3.2-Exp-Base scores higher on 14 and ERNIE 4.5 on 1. The widest gap is MMLU-Pro, where DeepSeek-V3.2-Exp-Base scores 63.3 against 16.
| Benchmark | DeepSeek-V3.2-Exp-Base | ERNIE 4.5 |
|---|---|---|
| arc_challenge | 95.5 | 40.6 |
| bbh | 88.7 | 30.4 |
| C-Eval | 91 | 40.7 |
| cmmlu | 88.9 | 39.8 |
| DROP | 86.6 | 28.6 |
| GPQA Diamond | 52 | 74 |
| GSM8K | 91.1 | 25.2 |
| hellaswag | 89.4 | 33 |
| HumanEval+ | 67.7 | 25 |
| MBPP+ | 69.8 | 40.2 |
| mmlu_redux | 90.4 | 43.2 |
| MMLU-Pro | 63.3 | 16 |
| PIQA | 85.1 | 55.2 |
| simpleqa | 27 | 1.8 |
| winogrande | 85.6 | 51.3 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.