VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, DeepSeek-V3.1-Base or ERNIE 4.5?
Across 14 shared benchmarks, DeepSeek-V3.1-Base scores higher on 13 and ERNIE 4.5 on 1. The widest gap is MMLU-Pro, where DeepSeek-V3.1-Base scores 58.8 against 16.

DeepSeek-V3.1-Base vs ERNIE 4.5

Across 14 shared benchmarks, DeepSeek-V3.1-Base scores higher on 13 and ERNIE 4.5 on 1. The widest gap is MMLU-Pro, where DeepSeek-V3.1-Base scores 58.8 against 16.

DeepSeekvsBaidu14 shared benchmarks131 head-to-head
BenchmarkDeepSeek-V3.1-BaseERNIE 4.5
arc_challenge95.640.6
bbh88.230.4
C-Eval9040.7
cmmlu88.839.8
DROP86.328.6
GPQA Diamond5174
GSM8K91.425.2
hellaswag89.233
HumanEval+64.625
MBPP+72.240.2
mmlu_redux9043.2
MMLU-Pro58.816
simpleqa26.31.8
winogrande85.951.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.