VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.7 Sonnet vs Gemini 2.5 Pro

AnthropicvsGoogle37 shared benchmarks1324 head-to-head
BenchmarkClaude 3.7 SonnetGemini 2.5 Pro
AA Agentic Index377.2
AA Intelligence27.627
AA-LCR62.366
AA-Omniscience-0.7-14.3
aider_polyglot64.982.2
AIME 202554.888.3
AIR-Bench 202481.873.6
anthropic_red_team99.799.5
arc_agi_113.641
ARC-AGI-204.9
arena_vision11951246
Artificial Analysis Coding Index36.446.7
bbq92.196.4
critpt0.92.6
Fortress3854.9
gdpval27.48.5
GPQA Diamond84.884.4
harmbench84.365.4
HLE9.722.5
hmmt_feb_202531.782.5
IFBench48.349
MMMU7583.9
MMMU-Pro60.174.9
multichallenge51.653.6
OmniScience Accuracy28.239
OmniScience Non-Hallucination6012.6
scicode40.343
SEAL VISTA48.250.8
simple_safety_tests10097
simplebench46.462.4
swe_bench_bash52.853.6
SWE-bench Verified70.367.2
Terminal-Bench35.225.3
Terminal-Bench Hard21.226.5
usamo_20253.624.4
xstest96.498.7
τ²-Bench Telecom (AA run)54.754.1

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.