VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

GPT-4o vs Qwen2.5 Instruct 72B

OpenAIvsAlibaba65 shared benchmarks4618 head-to-head
BenchmarkGPT-4oQwen2.5 Instruct 72B
AA Intelligence11.210
AA-LCR5520.3
AA-Omniscience-10.5-52.2
aider_edit_acc72.965.4
aider_polyglot30.77.6
AIME 20249.323.3
AIR-Bench 202462.459
anthropic_red_team99.199.6
Arena Hard92.481.2
Artificial Analysis Coding Index24.211.9
bbq95.195.4
C-Eval7689.2
Chinese SimpleQA (C-SimpleQA)64.652.2
CLUEWSC87.991.4
Codeforces75924.8
critpt00
DROP83.476.7
DROP (3-shot F1)83.776.7
DROP (F1)89.285
Fortress47.256.4
FRAMES (Acc.)80.569.8
GPQA Diamond70.149.1
GSM8K95.695.8
harmbench82.972.8
HLE5.34.2
humaneval90.686.6
HumanEval-Mul (Pass@1)80.577.3
IFBench3636.9
ifeval84.387.2
IFEval (avg)84.187.2
LiveCodeBench38.355.5
livecodebench_pass1cot34.231.1
LongBench v2 easy (w/ CoT)54.247.9
LongBench v2 easy (w/o CoT)57.442.7
LongBench v2 hard (w/ CoT)49.740.8
LongBench v2 hard (w/o CoT)45.641.8
LongBench v2 long (w/ CoT)43.539.8
LongBench v2 long (w/o CoT)40.244.4
LongBench v2 medium (w/ CoT)48.640.9
LongBench v2 medium (w/o CoT)52.438.1
LongBench v2 overall (w/ CoT)51.443.5
LongBench v2 overall (w/o CoT)50.142.1
LongBench v2 short (w/ CoT)59.648.9
LongBench v2 short (w/o CoT)53.345.6
longbench_v248.139.4
MATH85.388.4
math_500_em74.680
MBPP+ (EvalPlus-augmented)76.277
mgsm90.587.3
mmlu88.186.1
mmlu_redux8886.8
MMLU-Pro74.771.6
mmmlu81.474.8
narrativeqa80.474.5
naturalquestions_closedbook50.135.9
OmniScience Accuracy23.717.6
OmniScience Non-Hallucination62.115.3
OpenBookQA96.896.2
scicode33.426.7
simple_safety_tests98.5100
simpleqa39.410.3
SWE-bench Verified38.823.8
Terminal-Bench Hard8.34.5
xstest97.397.9
τ²-Bench Telecom (AA run)28.934.5

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.