Across 11 shared benchmarks, Qwen2.5 VL 32B Instruct scores higher on 9 and Qwen2 VL 72B Instruct on 2. The widest gap is mathvision, where Qwen2.5 VL 32B Instruct scores 40 against 25.9.
| Benchmark | Qwen2.5 VL 32B Instruct | Qwen2 VL 72B Instruct |
|---|---|---|
| CC-OCR | 77.1 | 68.7 |
| docvqa | 94.8 | 96.5 |
| InfoVQA | 83.4 | 84.5 |
| mathvision | 40 | 25.9 |
| MathVista | 74.7 | 70.5 |
| MMBench-Video | 1.9 | 1.7 |
| MMMU | 70 | 64.5 |
| MMMU-Pro | 49.5 | 46.2 |
| MMStar | 69.5 | 68.3 |
| ocrbench_v2 | 59.1 | 46.1 |
| Video-MME | 77.9 | 77.8 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.