Across 14 shared benchmarks, Qwen2 VL 72B Instruct scores higher on 8 and Qwen3 VL 4B (Reasoning) on 6. The widest gap is mathvision, where Qwen3 VL 4B (Reasoning) scores 60 against 25.9.
| Benchmark | Qwen2 VL 72B Instruct | Qwen3 VL 4B (Reasoning) |
|---|---|---|
| CC-OCR | 68.7 | 73.8 |
| chartqa | 88.3 | 84 |
| DocVQA_test | 96.5 | 94.2 |
| InfoVQA_test | 84.5 | 83 |
| mathvision | 25.9 | 60 |
| MathVista | 70.5 | 79.5 |
| MMMU (val) (Pass@1) | 64.5 | 70.8 |
| MMMU-Pro | 46.2 | 57 |
| MMStar | 68.3 | 73.2 |
| MV-Bench | 73.6 | 69.3 |
| OCRBench | 877 | 81.6 |
| RealWorldQA | 77.8 | 73.2 |
| TextVQA-val | 85.5 | 80.5 |
| Video-MME | 77.8 | 59.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.