Across 12 shared benchmarks, Qwen2 VL 72B Instruct scores higher on 1 and Qwen3 VL 235B A22B Reasoning on 10, with 1 level. The widest gap is mathvision, where Qwen3 VL 235B A22B Reasoning scores 74.6 against 25.9.
| Benchmark | Qwen2 VL 72B Instruct | Qwen3 VL 235B A22B Reasoning |
|---|---|---|
| CC-OCR | 68.7 | 81.5 |
| DocVQA_test | 96.5 | 96.5 |
| InfoVQA_test | 84.5 | 89.5 |
| mathvision | 25.9 | 74.6 |
| MathVista | 70.5 | 85.8 |
| MMMU | 64.5 | 78.7 |
| MMMU (val) (Pass@1) | 64.5 | 80.6 |
| MMMU-Pro | 46.2 | 69.3 |
| MMStar | 68.3 | 78.7 |
| OCRBench | 877 | 87.5 |
| RealWorldQA | 77.8 | 81.3 |
| Video-MME | 77.8 | 79 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.