VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

o1 vs o4-mini

OpenAIvsOpenAI50 shared benchmarks2624 head-to-head
Benchmarko1o4-mini
AA Agentic Index31.136.1
AA Intelligence2426.1
AA-LCR63.360
AA-Omniscience-10.6-35.7
aider_polyglot61.772
AIME 202581.792.7
AIR-Bench 20248078.5
anthropic_red_team98.398.2
arena_vision11931201
Artificial Analysis Coding Index39.725.6
bbq97.394
BBQ Accuracy on Ambiguous Questions10.8
BBQ Accuracy on Unambiguous Questions0.90.9
BBQ P(not stereotyping | ambiguous question, not unknown)0.10.3
critpt0.30.6
Fortress19.421.5
gdpval11.525.4
GPQA Diamond75.781.4
GSM8K97.10
harmbench96.397
HLE7.716.5
IFBench70.368.7
MathVista71.884.3
MMLU Language (0-shot) - Average0.90.9
MMLU Language (0-shot) - Chinese (Simplified)0.90.9
MMLU Language (0-shot) - French0.90.9
MMLU Language (0-shot) - Hindi0.90.9
MMLU Language (0-shot) - Indonesian0.90.9
MMLU Language (0-shot) - Italian0.90.9
MMLU Language (0-shot) - Japanese0.90.9
MMLU Language (0-shot) - Korean0.90.9
MMLU Language (0-shot) - Portuguese (Brazil)0.90.9
MMLU Language (0-shot) - Swahili0.90.8
MMLU Language (0-shot) - Yoruba0.80.7
MMMU77.681.6
OmniScience Accuracy34.724.8
OmniScience Non-Hallucination30.719.5
PersonQA hallucination rate0.20.5
scicode35.846.5
SEAL VISTA45.351.8
simple_safety_tests99100
simplebench41.738.7
SimpleQA hallucination rate0.40.8
strongreject9796
SWE-bench Verified48.968.1
TAU-bench (airline)5049.2
TAU-bench (retail)70.871.8
Terminal-Bench Hard12.915.2
xstest9797.4
τ²-Bench Telecom (AA run)62.655.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.