VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Claude Opus 4.5 vs o3

AnthropicvsOpenAI40 shared benchmarks3010 head-to-head
BenchmarkClaude Opus 4.5o3
AA Agentic Index59.636.1
AA Intelligence41.931.1
AA-LCR7673.3
AA-Omniscience14-15.3
AIME 202592.889.2
arc_agi_18260.8
ARC-AGI-237.66.5
Artificial Analysis Coding Index47.838.4
browsecomp3749.7
critpt4.61.1
ERQA46.864
Fortress13.616
frontiermath_tier_44.22.1
gdpval47.312.8
GPQA Diamond8783.3
HLE30.820.3
hmmt_feb_202592.977.5
IFBench5871.4
LiveCodeBench8784.7
longbench_v264.458.8
MathVista80.286.8
metr_hcast73.464.7
metr_re_bench4015
metr_swaa99.599.8
MMLU-Pro9085
MMMU80.782.9
MMMU-Pro7476.4
multichallenge5960.4
MultiNRC48.645.5
OmniScience Accuracy46.638.6
OmniScience Non-Hallucination3912.9
scicode5041
simplebench6253.1
simpleqa_verified41.853
swe_bench_bash76.858.4
SWE-bench Verified81.569.1
Tau2 retail88.980.2
Terminal-Bench Hard4737.1
VideoMMMU84.483.3
τ²-Bench Telecom (AA run)89.580.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.