VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-R1 vs o3

DeepSeekvsOpenAI49 shared benchmarks840 head-to-head
BenchmarkDeepSeek-R1o3
AA Agentic Index3.136.1
AA Intelligence18.631.1
AA-LCR5673.3
AA-Omniscience-31.3-15.3
aider_polyglot71.681.3
AIME 202491.491.6
AIME 202587.589.2
AIR-Bench 202452.984.5
anthropic_red_team99.198.3
arc_agi_121.260.8
ARC-AGI-21.36.5
Artificial Analysis Coding Index24.638.4
bbq96.697.9
browsecomp8.949.7
critpt0.61.1
Fortress74.416
FullStackBench70.169.3
gdpval1.512.8
GPQA Diamond8183.3
harmbench54.698.4
HLE17.720.3
hmmt_feb_202576.777.5
IFBench3971.4
imo_20256.816.7
LiveCodeBench84.484.7
LiveCodeBench (24/8~25/5)73.175.8
livecodebench_easy99.299.1
livecodebench_hard63.666
livecodebench_medium90.989.8
longbench_v258.358.8
math_500_em9898.1
MMLU-Pro8585
multichallenge4560.4
MultiNRC27.645.5
OmniScience Accuracy30.738.6
OmniScience Non-Hallucination10.512.9
OpenAI-MRCR (128k)51.556.5
scicode35.741
simple_safety_tests98.399
simplebench40.853.1
simpleqa92.349.4
simpleqa_verified27.453
SWE-bench Verified57.669.1
TAU-bench (airline)53.552
TAU-bench (retail)63.973.9
Terminal-Bench Hard6.137.1
xstest98.897.3
ZebraLogic95.195.8
τ²-Bench Telecom (AA run)11.480.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.