VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.7 Sonnet vs DeepSeek-R1

AnthropicvsDeepSeek37 shared benchmarks2512 head-to-head
BenchmarkClaude 3.7 SonnetDeepSeek-R1
AA Agentic Index373.1
AA Intelligence27.618.6
AA-LCR62.356
AA-Omniscience-0.7-31.3
aider_polyglot64.971.6
AIME 202554.887.5
AIR-Bench 202481.852.9
anthropic_red_team99.799.1
arc_agi_113.621.2
ARC-AGI-201.3
Artificial Analysis Coding Index36.424.6
bbq92.196.6
critpt0.90.6
Fortress3874.4
gdpval27.41.5
GPQA Diamond84.881
harmbench84.354.6
HLE9.717.7
hmmt_feb_202531.776.7
IFBench48.339
ifeval93.283.3
math_500_em96.298
multichallenge51.645
MultiNRC27.827.6
OmniScience Accuracy28.230.7
OmniScience Non-Hallucination6010.5
scicode40.335.7
simple_safety_tests10098.3
simplebench46.440.8
SWE-bench Verified70.357.6
TAU-bench (airline)58.453.5
TAU-bench (retail)81.263.9
Terminal-Bench35.25.7
Terminal-Bench Hard21.26.1
usamo_20253.630.1
xstest96.498.8
τ²-Bench Telecom (AA run)54.711.4

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.