VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.5 Sonnet vs DeepSeek-R1

AnthropicvsDeepSeek42 shared benchmarks933 head-to-head
BenchmarkClaude 3.5 SonnetDeepSeek-R1
AA Intelligence1018.6
aider_polyglot45.371.6
AIME 20241691.4
AIME 20257.487.5
AIR-Bench 202485.952.9
AlpacaEval2.0 (LC-winrate)5287.6
anthropic_red_team99.899.1
Arena-Hard (GPT-4-1106 judge)85.292.3
ArenaHard (GPT-4-1106)85.292.3
Artificial Analysis Coding Index30.224.6
bbq94.996.6
C-Eval76.791.8
Chinese SimpleQA (C-SimpleQA)56.863.7
CLUEWSC85.492.8
CNMO 202413.178.8
Codeforces7172029
Codeforces (Percentile)20.396.3
Codeforces (Rating)7172029
DROP (3-shot F1)88.392.2
Fortress1374.4
FRAMES (Acc.)72.583
GPQA Diamond67.281
harmbench98.154.6
HLE3.917.7
hmmt_feb_20251.776.7
ifeval90.183.3
LiveCodeBench32.884.4
livecodebench_pass1cot36.365.9
longbench_v24158.3
MATH81.397.3
math_500_em78.398
mmlu88.392.9
mmlu_redux88.993.4
MMLU-Pro7885
scicode36.635.7
simple_safety_tests10098.3
simplebench41.440.8
simpleqa28.492.3
SWE-bench Verified50.857.6
TAU-bench (airline)4653.5
TAU-bench (retail)69.263.9
xstest95.698.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.