Across 15 shared benchmarks, Qwen2.5 VL 32B Instruct scores higher on 2 and Qwen3.5 35B A3B on 13. The widest gap is OSWorld-Verified, where Qwen3.5 35B A3B scores 54.5 against 5.9.
| Benchmark | Qwen2.5 VL 32B Instruct | Qwen3.5 35B A3B |
|---|---|---|
| CC-OCR | 77.1 | 80.7 |
| GPQA Diamond | 46 | 84.5 |
| humaneval | 91.5 | 66.5 |
| LVBench | 49 | 71.4 |
| mathvision | 40 | 83.9 |
| MathVista | 74.7 | 86.2 |
| MBPP | 84 | 70.8 |
| mmlu | 78.4 | 81.1 |
| MMLU-Pro | 68.8 | 85.3 |
| MMMU | 70 | 81.4 |
| MMMU-Pro | 49.5 | 75.1 |
| MMStar | 69.5 | 81.9 |
| ocrbench_v2 | 59.1 | 65.3 |
| OSWorld-Verified | 5.9 | 54.5 |
| screenspot_pro_no_tools | 39.4 | 68.6 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.