VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-4o vs o1

OpenAIvsOpenAI49 shared benchmarks642 head-to-head
BenchmarkGPT-4oo1
AA Agentic Index9.731.1
AA Intelligence11.224
AA-LCR5563.3
AA-Omniscience-10.5-10.6
aider_polyglot30.761.7
AIME 20249.379.2
AIME 202511.781.7
AIR-Bench 202462.480
anthropic_red_team99.198.3
arena_vision11621193
Artificial Analysis Coding Index24.239.7
bbq95.197.3
Codeforces7592061
Codeforces (Percentile)23.696.6
Codeforces (Rating)7592061
critpt00.3
DROP (3-shot F1)83.790.2
Fortress47.219.4
gdpval011.5
GPQA Diamond70.175.7
GSM8K95.697.1
harmbench82.996.3
HLE5.37.7
humaneval90.688.1
IFBench3670.3
livecodebench_pass1cot34.263.4
MATH85.396.4
math_500_em74.696.4
MathVista63.871.8
metr_hcast2542.6
metr_re_bench00
metr_swaa99.2100
mgsm90.590.8
mmlu88.191.8
mmmlu81.487.7
MMMU72.277.6
OmniScience Accuracy23.734.7
OmniScience Non-Hallucination62.130.7
scicode33.435.8
SEAL VISTA34.945.3
simple_safety_tests98.599
simplebench17.841.7
simpleqa39.447
SWE-bench Verified38.848.9
TAU-bench (airline)42.850
TAU-bench (retail)60.370.8
Terminal-Bench Hard8.312.9
xstest97.397
τ²-Bench Telecom (AA run)28.962.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.