VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, GPT-4o or Qwen2 VL 72B Instruct?
Across 24 shared benchmarks, GPT-4o scores higher on 6 and Qwen2 VL 72B Instruct on 18. The widest gap is MMMU-Pro, where GPT-4o scores 59.9 against 46.2.

GPT-4o vs Qwen2 VL 72B Instruct

Across 24 shared benchmarks, GPT-4o scores higher on 6 and Qwen2 VL 72B Instruct on 18. The widest gap is MMMU-Pro, where GPT-4o scores 59.9 against 46.2.

OpenAIvsAlibaba24 shared benchmarks618 head-to-head
BenchmarkGPT-4oQwen2 VL 72B Instruct
Avg.27.830.9
chartqa88.188.3
docvqa92.896.5
DocVQA_test92.896.5
EgoSchema72.277.9
EgoSchema_test72.277.9
HallBench_avg5558.1
mathvision30.425.9
MathVista63.870.5
MMBench-CN_test82.186.6
MMBench-EN_test83.486.5
MMBench-V1.1_test82.285.9
MME_sum23292483
MMMU72.264.5
MMMU (val) (Pass@1)69.164.5
MMMU-Pro59.946.2
MMStar63.968.3
MMT-Bench_test65.571.7
OCRBench806877
RealWorldQA75.477.8
REVERIE_valid-unseen SR31.631
VCR_en easy91.591.9
Video-MME77.277.8
Video-MME (wo subs)71.971.2

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.