VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.5 Sonnet vs DeepSeek-V3

AnthropicvsDeepSeek56 shared benchmarks1838 head-to-head
BenchmarkClaude 3.5 SonnetDeepSeek-V3
AA Intelligence1015.4
aider_edit_acc84.279.7
aider_polyglot45.355.1
AIME 20241659.4
AIME 20257.451.3
AIR-Bench 202485.940.8
AlpacaEval2.0 (LC-winrate)5270
anthropic_red_team99.897.1
Arena Hard87.691.4
Arena-Hard (GPT-4-1106 judge)85.285.5
ArenaHard (GPT-4-1106)85.285.5
Artificial Analysis Coding Index30.223
bbh93.187.5
bbq94.996.7
C-Eval76.790.1
Chinese SimpleQA (C-SimpleQA)56.868
CLUEWSC85.490.9
CNMO 202413.174.7
Codeforces7171134
Codeforces (Percentile)20.358.7
Codeforces (Rating)7171134
DROP87.191.6
DROP (3-shot F1)88.391.6
DROP (F1)88.891
FRAMES (Acc.)72.573.3
GPQA Diamond67.268.4
GSM8K96.996.7
harmbench98.149.7
HLE3.95.2
hmmt_feb_20251.729.2
humaneval93.792.1
HumanEval-Mul (Pass@1)81.782.6
ifeval90.187.3
IFEval (avg)90.187.3
LiveCodeBench32.849.6
LiveCodeBench v637.246.9
livecodebench_pass1cot36.340.5
LongBench v2 overall (w/o CoT)4148.7
longbench_v24148.7
MATH81.391.2
math_500_em78.394
MBPP+ (EvalPlus-augmented)75.178.8
mgsm91.679.8
mmlu88.389.4
mmlu_redux88.990.5
MMLU-Pro7881.2
narrativeqa74.679.6
naturalquestions_closedbook50.246.7
OpenBookQA97.295.4
scicode36.635.8
simple_safety_tests10095.3
simplebench41.427.2
simpleqa28.427.7
supergpqa48.253.7
SWE-bench Verified50.842
xstest95.697.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.