VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, GPT-4o or Qwen2.5 VL 32B Instruct?
Across 14 shared benchmarks, GPT-4o scores higher on 7 and Qwen2.5 VL 32B Instruct on 7. The widest gap is GPQA Diamond, where GPT-4o scores 70.1 against 46.

GPT-4o vs Qwen2.5 VL 32B Instruct

Across 14 shared benchmarks, GPT-4o scores higher on 7 and Qwen2.5 VL 32B Instruct on 7. The widest gap is GPQA Diamond, where GPT-4o scores 70.1 against 46.

OpenAIvsAlibaba14 shared benchmarks77 head-to-head
BenchmarkGPT-4oQwen2.5 VL 32B Instruct
arena_vision11621153
docvqa92.894.8
GPQA Diamond70.146
humaneval90.691.5
MATH85.382.2
mathvision30.440
MathVista63.874.7
mmlu88.178.4
MMLU-Pro74.768.8
MMMU72.270
MMMU-Pro59.949.5
MMStar63.969.5
OSWorld-Verified55.9
Video-MME77.277.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.