VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-R1 vs o1

DeepSeekvsOpenAI38 shared benchmarks1721 head-to-head
BenchmarkDeepSeek-R1o1
AA Agentic Index3.131.1
AA Intelligence18.624
AA-LCR5663.3
AA-Omniscience-31.3-10.6
aider_polyglot71.661.7
AIME 202491.479.2
AIME 202587.581.7
AIR-Bench 202452.980
anthropic_red_team99.198.3
Artificial Analysis Coding Index24.639.7
bbq96.697.3
Codeforces20292061
Codeforces (Percentile)96.396.6
Codeforces (Rating)20292061
critpt0.60.3
DROP (3-shot F1)92.290.2
Fortress74.419.4
gdpval1.511.5
GPQA Diamond8175.7
harmbench54.696.3
HLE17.77.7
IFBench3970.3
livecodebench_pass1cot65.963.4
MATH97.396.4
math_500_em9896.4
mmlu92.991.8
OmniScience Accuracy30.734.7
OmniScience Non-Hallucination10.530.7
scicode35.735.8
simple_safety_tests98.399
simplebench40.841.7
simpleqa92.347
SWE-bench Verified57.648.9
TAU-bench (airline)53.550
TAU-bench (retail)63.970.8
Terminal-Bench Hard6.112.9
xstest98.897
τ²-Bench Telecom (AA run)11.462.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.