VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

o1 vs Qwen2.5 Instruct 72B

OpenAIvsAlibaba35 shared benchmarks304 head-to-head
Benchmarko1Qwen2.5 Instruct 72B
AA Intelligence2410
AA-LCR63.320.3
AA-Omniscience-10.6-52.2
aider_polyglot61.77.6
AIME 202479.223.3
AIR-Bench 20248059
anthropic_red_team98.399.6
Artificial Analysis Coding Index39.711.9
bbq97.395.4
Codeforces206124.8
critpt0.30
DROP (3-shot F1)90.276.7
Fortress19.456.4
GPQA Diamond75.749.1
GSM8K97.195.8
harmbench96.372.8
HLE7.74.2
humaneval88.186.6
IFBench70.336.9
livebench52.352.3
livecodebench_pass1cot63.431.1
MATH96.488.4
math_500_em96.480
mgsm90.887.3
mmlu91.886.1
mmmlu87.774.8
OmniScience Accuracy34.717.6
OmniScience Non-Hallucination30.715.3
scicode35.826.7
simple_safety_tests99100
simpleqa4710.3
SWE-bench Verified48.923.8
Terminal-Bench Hard12.94.5
xstest9797.9
τ²-Bench Telecom (AA run)62.634.5

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.