VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

o1 vs o3

OpenAIvsOpenAI53 shared benchmarks743 head-to-head
Benchmarko1o3
AA Agentic Index31.136.1
AA Intelligence2431.1
AA-LCR63.373.3
AA-Omniscience-10.6-15.3
aider_polyglot61.781.3
AIME 202479.291.6
AIME 202581.789.2
AIR-Bench 20248084.5
anthropic_red_team98.398.3
arena_vision11931217
Artificial Analysis Coding Index39.738.4
bbq97.397.9
BBQ Accuracy on Ambiguous Questions10.9
BBQ Accuracy on Unambiguous Questions0.90.9
BBQ P(not stereotyping | ambiguous question, not unknown)0.10.3
critpt0.31.1
Fortress19.416
gdpval11.512.8
GPQA Diamond75.783.3
harmbench96.398.4
HLE7.720.3
IFBench70.371.4
math_500_em96.498.1
MathVista71.886.8
metr_hcast42.664.7
metr_re_bench015
metr_swaa10099.8
MMLU Language (0-shot) - Average0.90.9
MMLU Language (0-shot) - Bengali0.90.9
MMLU Language (0-shot) - Chinese (Simplified)0.90.9
MMLU Language (0-shot) - French0.90.9
MMLU Language (0-shot) - Hindi0.90.9
MMLU Language (0-shot) - Indonesian0.90.9
MMLU Language (0-shot) - Italian0.90.9
MMLU Language (0-shot) - Korean0.90.9
MMLU Language (0-shot) - Spanish0.90.9
MMMU77.682.9
OmniScience Accuracy34.738.6
OmniScience Non-Hallucination30.712.9
PersonQA hallucination rate0.20.3
scicode35.841
simple_safety_tests9999
simplebench41.753.1
simpleqa4749.4
SimpleQA accuracy0.50.5
SimpleQA hallucination rate0.40.5
strongreject9797
SWE-bench Verified48.969.1
TAU-bench (airline)5052
TAU-bench (retail)70.873.9
Terminal-Bench Hard12.937.1
xstest9797.3
τ²-Bench Telecom (AA run)62.680.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.