VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-4o vs o3

OpenAIvsOpenAI54 shared benchmarks548 head-to-head
BenchmarkGPT-4oo3
AA Agentic Index9.736.1
AA Intelligence11.231.1
AA-LCR5573.3
AA-Omniscience-10.5-15.3
aider_polyglot30.781.3
AIME 20249.391.6
AIME 202511.789.2
AIR-Bench 202462.484.5
anthropic_red_team99.198.3
arc_agi_14.560.8
ARC-AGI-206.5
arena_vision11621217
Artificial Analysis Coding Index24.238.4
bbq95.197.9
critpt01.1
ERQA35.264
Fortress47.216
gdpval012.8
GPQA Diamond70.183.3
harmbench82.998.4
HLE5.320.3
IFBench3671.4
LiveCodeBench38.384.7
livecodebench_easy82.599.1
livecodebench_hard4.566
livecodebench_medium32.189.8
longbench_v248.158.8
math_500_em74.698.1
MathVista63.886.8
metr_hcast2564.7
metr_re_bench015
metr_swaa99.299.8
MMLU-Pro74.785
MMMU72.282.9
MMMU-Pro59.976.4
multichallenge40.360.4
MultiNRC12.445.5
OmniScience Accuracy23.738.6
OmniScience Non-Hallucination62.112.9
PropensityBench46.110.5
scicode33.441
simple_safety_tests98.599
simplebench17.853.1
simpleqa39.449.4
swe_bench_bash21.658.4
SWE-bench Verified38.869.1
TAU-bench (airline)42.852
TAU-bench (retail)60.373.9
Tau2 airline45.564.8
Tau2 retail63.480.2
Terminal-Bench Hard8.337.1
VideoMMMU61.283.3
xstest97.397.3
τ²-Bench Telecom (AA run)28.980.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.