VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-R1 vs Gemini 2.5 Pro

DeepSeekvsGoogle50 shared benchmarks1238 head-to-head
BenchmarkDeepSeek-R1Gemini 2.5 Pro
AA Agentic Index3.17.2
AA Intelligence18.627
AA-LCR5666
AA-Omniscience-31.3-14.3
aider_polyglot71.682.2
AIME 202587.588.3
AIR-Bench 202452.973.6
anthropic_red_team99.199.5
arc_agi_121.241
ARC-AGI-21.34.9
Artificial Analysis Coding Index24.646.7
bbq96.696.4
browsecomp8.99.9
browsecomp_zh35.732.2
critpt0.62.6
Fortress74.454.9
gdpval1.58.5
GPQA Diamond8184.4
harmbench54.665.4
HLE17.722.5
HMMT 202579.465.7
hmmt_feb_202576.782.5
IFBench3949
imo_20256.831.6
LiveCodeBench84.482.7
livecodebench_easy99.298.8
livecodebench_hard63.659.4
livecodebench_medium90.990.6
MMLU-Pro8586
multichallenge4553.6
OmniScience Accuracy30.739
OmniScience Non-Hallucination10.512.6
scicode35.743
simple_safety_tests98.397
simplebench40.862.4
simpleqa92.350.8
simpleqa_verified27.456
SWE-bench Verified57.667.2
TauBench V3 - Banking6.49.7
Terminal-Bench5.725.3
Terminal-Bench 2.119.128.5
Terminal-Bench Hard6.126.5
usamo_202530.124.4
vectara_answer_rate9799.1
vectara_avg_summary_length93.5106.4
vectara_factual_consistency88.793
vectara_hallucination_rate11.37
xstest98.898.7
τ²-Bench Telecom (AA run)11.454.1
τ³-Bench6.49.3

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.