Across 17 shared benchmarks, Qwen2.5 VL 32B scores higher on 0 and Qwen2.5 VL 32B Instruct on 2, with 15 level. The widest gap is mathvision, where Qwen2.5 VL 32B Instruct scores 40 against 38.4.
| Benchmark | Qwen2.5 VL 32B | Qwen2.5 VL 32B Instruct |
|---|---|---|
| CC-OCR | 77.1 | 77.1 |
| docvqa | 94.8 | 94.8 |
| GPQA Diamond | 46 | 46 |
| humaneval | 91.5 | 91.5 |
| LVBench | 49 | 49 |
| MATH | 82.2 | 82.2 |
| mathvision | 38.4 | 40 |
| MathVista | 74.7 | 74.7 |
| MBPP | 0.8 | 84 |
| mmlu | 78.4 | 78.4 |
| MMLU-Pro | 68.8 | 68.8 |
| MMMU | 70 | 70 |
| MMMU-Pro | 49.5 | 49.5 |
| MMStar | 69.5 | 69.5 |
| OSWorld-Verified | 5.9 | 5.9 |
| screenspot_pro_no_tools | 39.4 | 39.4 |
| Video-MME | 77.9 | 77.9 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.