VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude 3.5 Sonnet vs Claude 4 Opus

AnthropicvsAnthropic35 shared benchmarks529 head-to-head
BenchmarkClaude 3.5 SonnetClaude 4 Opus
AA Intelligence1031.7
aider_polyglot45.370.7
AIME 20241676
AIME 20257.475.5
AIR-Bench 202485.985.7
anthropic_red_team99.898.9
arena_vision11461207
Artificial Analysis Coding Index30.234
bbq94.999.3
CNMO 202413.157.6
Fortress1327.6
frontiermath_tier_404.2
GPQA Diamond67.279.6
harmbench98.191.7
HLE3.912.3
hmmt_feb_20251.760
ifeval90.187.4
LiveCodeBench32.870.4
LiveCodeBench v637.247.4
longbench_v24155.6
math_500_em78.398.2
mmlu88.392.9
mmlu_redux88.994.2
MMLU-Pro7886.6
MMMU7273.7
scicode36.640.9
SEAL VISTA38.747
simple_safety_tests100100
simplebench41.458.8
simpleqa28.422.8
supergpqa48.256.5
SWE-bench Verified50.879.4
TAU-bench (airline)4659.6
TAU-bench (retail)69.281.4
xstest95.697

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.