Across 14 shared benchmarks, GPT-4o scores higher on 7 and Qwen2.5 VL 32B Instruct on 7. The widest gap is GPQA Diamond, where GPT-4o scores 70.1 against 46.
| Benchmark | GPT-4o | Qwen2.5 VL 32B Instruct |
|---|---|---|
| arena_vision | 1162 | 1153 |
| docvqa | 92.8 | 94.8 |
| GPQA Diamond | 70.1 | 46 |
| humaneval | 90.6 | 91.5 |
| MATH | 85.3 | 82.2 |
| mathvision | 30.4 | 40 |
| MathVista | 63.8 | 74.7 |
| mmlu | 88.1 | 78.4 |
| MMLU-Pro | 74.7 | 68.8 |
| MMMU | 72.2 | 70 |
| MMMU-Pro | 59.9 | 49.5 |
| MMStar | 63.9 | 69.5 |
| OSWorld-Verified | 5 | 5.9 |
| Video-MME | 77.2 | 77.9 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.