VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Sonnet 4 vs o1

AnthropicvsOpenAI37 shared benchmarks2710 head-to-head
BenchmarkClaude Sonnet 4o1
AA Agentic Index39.231.1
AA Intelligence29.824
AA-LCR69.763.3
AA-Omniscience0.2-10.6
aider_polyglot70.761.7
AIME 202443.479.2
AIME 20257481.7
AIR-Bench 202488.380
anthropic_red_team99.298.3
arena_vision12071193
Artificial Analysis Coding Index37.639.7
bbq97.997.3
critpt1.10.3
Fortress24.419.4
gdpval31.211.5
GPQA Diamond7875.7
harmbench9896.3
HLE10.77.7
IFBench5570.3
livebench74.852.3
math_500_em9496.4
mmlu91.591.8
mmmlu86.587.7
MMMU74.477.6
OmniScience Accuracy22.734.7
OmniScience Non-Hallucination70.930.7
scicode4035.8
SEAL VISTA45.545.3
simple_safety_tests10099
simplebench45.541.7
simpleqa15.947
SWE-bench Verified80.248.9
TAU-bench (airline)6050
TAU-bench (retail)80.570.8
Terminal-Bench Hard31.112.9
xstest97.297
τ²-Bench Telecom (AA run)64.662.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.