VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-4o vs o4-mini

OpenAIvsOpenAI49 shared benchmarks940 head-to-head
BenchmarkGPT-4oo4-mini
AA Agentic Index9.736.1
AA Intelligence11.226.1
AA-LCR5560
AA-Omniscience-10.5-35.7
aider_polyglot30.772
AIME 202511.792.7
AIR-Bench 202462.478.5
anthropic_red_team99.198.2
arc_agi_14.558.7
ARC-AGI-206.1
arena_vision11621201
Artificial Analysis Coding Index24.225.6
bbq95.194
critpt00.6
Fortress47.221.5
gdpval025.4
GPQA Diamond70.181.4
GSM8K95.60
harmbench82.997
HLE5.316.5
IFBench3668.7
LiveCodeBench38.384.5
livecodebench_easy82.598.8
livecodebench_hard4.562.9
livecodebench_medium32.192.2
MathVista63.884.3
mmlu_prox61.169.3
MMMU72.281.6
MMMU-Pro59.969.2
multichallenge40.344.9
MultiNRC12.422.2
OmniScience Accuracy23.724.8
OmniScience Non-Hallucination62.119.5
PropensityBench46.115.8
scicode33.446.5
SEAL VISTA34.951.8
simple_safety_tests98.5100
simplebench17.838.7
swe_bench_bash21.645
SWE-bench Verified38.868.1
TAU-bench (airline)42.849.2
TAU-bench (retail)60.371.8
Terminal-Bench Hard8.315.2
vectara_answer_rate93.899.2
vectara_avg_summary_length86.6130.9
vectara_factual_consistency90.481.4
vectara_hallucination_rate9.618.6
xstest97.397.4
τ²-Bench Telecom (AA run)28.955.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.