VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V3 vs GPT-4o

DeepSeekvsOpenAI76 shared benchmarks5024 head-to-head
BenchmarkDeepSeek-V3GPT-4o
AA Agentic Index1.69.7
AA Intelligence15.411.2
AA-LCR41.355
AA-Omniscience-37.6-10.5
aider_edit_acc79.772.9
aider_polyglot55.130.7
AIME 202459.49.3
AIME 202551.311.7
AIR-Bench 202440.862.4
AlpacaEval2.0 (LC-winrate)7051.1
anthropic_red_team97.199.1
Arena Hard91.492.4
Arena-Hard (GPT-4-1106 judge)85.580.4
ArenaHard (GPT-4-1106)85.580.4
Artificial Analysis Coding Index2324.2
bbq96.795.1
C-Eval90.176
Chinese SimpleQA (C-SimpleQA)6864.6
CLUEWSC90.987.9
CNMO 202474.710.8
Codeforces1134759
Codeforces (Percentile)58.723.6
Codeforces (Rating)1134759
critpt00
DROP91.683.4
DROP (3-shot F1)91.683.7
DROP (F1)9189.2
FRAMES (Acc.)73.380.5
gdpval00
GPQA Diamond68.470.1
GSM8K96.795.6
harmbench49.782.9
HLE5.25.3
humaneval92.190.6
HumanEval-Mul (Pass@1)82.680.5
IFBench4136
ifeval87.384.3
IFEval (avg)87.384.1
LiveCodeBench49.638.3
LiveCodeBench v646.930.9
livecodebench_easy83.382.5
livecodebench_hard14.24.5
livecodebench_medium53.532.1
livecodebench_pass1cot40.534.2
LongBench v2 overall (w/o CoT)48.750.1
longbench_v248.748.1
MATH91.285.3
math_500_em9474.6
MBPP+ (EvalPlus-augmented)78.876.2
mgsm79.890.5
mmlu89.488.1
mmlu_prox70.561.1
mmlu_redux90.588
MMLU-Pro81.274.7
mmmlu79.481.4
multichallenge31.440.3
narrativeqa79.680.4
naturalquestions_closedbook46.750.1
OmniScience Accuracy25.423.7
OmniScience Non-Hallucination23.362.1
OpenBookQA95.496.8
scicode35.833.4
simple_safety_tests95.398.5
simplebench27.217.8
simpleqa27.739.4
supergpqa53.742.4
SWE-bench Verified4238.8
Tau2 airline3945.5
Tau2 retail69.163.4
Terminal-Bench Hard15.28.3
vectara_answer_rate97.593.8
vectara_avg_summary_length81.786.6
vectara_factual_consistency93.990.4
vectara_hallucination_rate6.19.6
xstest97.197.3
τ²-Bench Telecom (AA run)47.128.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.