VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

o3 vs Qwen2.5 Instruct 72B

OpenAIvsAlibaba28 shared benchmarks235 head-to-head
Benchmarko3Qwen2.5 Instruct 72B
AA Intelligence31.110
AA-LCR73.320.3
AA-Omniscience-15.3-52.2
aider_polyglot81.37.6
AIME 202491.623.3
AIR-Bench 202484.559
anthropic_red_team98.399.6
Artificial Analysis Coding Index38.411.9
bbq97.995.4
critpt1.10
Fortress1656.4
GPQA Diamond83.349.1
harmbench98.472.8
HLE20.34.2
IFBench71.436.9
LiveCodeBench84.755.5
longbench_v258.839.4
math_500_em98.180
MMLU-Pro8571.6
OmniScience Accuracy38.617.6
OmniScience Non-Hallucination12.915.3
scicode4126.7
simple_safety_tests99100
simpleqa49.410.3
SWE-bench Verified69.123.8
Terminal-Bench Hard37.14.5
xstest97.397.9
τ²-Bench Telecom (AA run)80.734.5

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.