Across 13 shared benchmarks, Qwen2 VL 72B Instruct scores higher on 5 and Qwen3 VL Thinking (8B) on 8. The widest gap is mathvision, where Qwen3 VL Thinking (8B) scores 62.7 against 25.9.
| Benchmark | Qwen2 VL 72B Instruct | Qwen3 VL Thinking (8B) |
|---|---|---|
| CC-OCR | 68.7 | 76.3 |
| DocVQA_test | 96.5 | 95.3 |
| InfoVQA_test | 84.5 | 86 |
| mathvision | 25.9 | 62.7 |
| MathVista | 70.5 | 81.4 |
| MMMU | 64.5 | 73.5 |
| MMMU (val) (Pass@1) | 64.5 | 74.1 |
| MMMU-Pro | 46.2 | 60.4 |
| MMStar | 68.3 | 75.3 |
| MV-Bench | 73.6 | 69 |
| OCRBench | 877 | 82.8 |
| RealWorldQA | 77.8 | 73.5 |
| Video-MME | 77.8 | 71.8 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.