Across 40 shared benchmarks, Qwen3.5 122B A10B scores higher on 38 and Qwen3 VL 4B (Reasoning) on 2. The widest gap is Artificial Analysis Coding Index, where Qwen3.5 122B A10B scores 45.7 against 6.7.
| Benchmark | Qwen3.5 122B A10B | Qwen3 VL 4B (Reasoning) |
|---|---|---|
| AA Agentic Index | 21.3 | 14.4 |
| AA Intelligence | 32.8 | 7.7 |
| AA-LCR | 70.3 | 23 |
| AA-Omniscience | -41.5 | -68.4 |
| AI2D | 93.3 | 84.9 |
| Artificial Analysis Coding Index | 45.7 | 6.7 |
| CC-OCR | 81.8 | 73.8 |
| critpt | 0.9 | 0 |
| EmbSpatial-Bench | 0.8 | 80.7 |
| ERQA | 62 | 47.3 |
| gdpval | 24.3 | 13.8 |
| GPQA Diamond | 86.6 | 64.1 |
| HallusionBench | 67.6 | 64.1 |
| HLE | 47.5 | 4.6 |
| HMMT 2025 | 90.3 | 53.1 |
| IFBench | 76.1 | 36.6 |
| ifeval | 93.4 | 82.6 |
| include | 82.8 | 64.6 |
| LiveCodeBench v6 | 78.9 | 51.3 |
| LVBench | 74.4 | 53.5 |
| mathvision | 86.2 | 60 |
| MathVista | 87.4 | 79.5 |
| mmlu_prox | 82.2 | 65 |
| mmlu_redux | 94 | 86 |
| MMLU-Pro | 86.7 | 73.6 |
| MMMU-Pro | 76.9 | 57 |
| MMStar | 82.9 | 73.2 |
| MV-Bench | 76.6 | 69.3 |
| OCRBench | 92.1 | 81.6 |
| OmniScience Accuracy | 24.4 | 12.2 |
| OmniScience Non-Hallucination | 12.9 | 8.2 |
| OSWorld-Verified | 58 | 31.4 |
| RealWorldQA | 85.1 | 73.2 |
| RefSpatial-Bench | 0.7 | 45.3 |
| scicode | 42 | 17.1 |
| screenspot_pro_no_tools | 70.4 | 49.2 |
| supergpqa | 67.1 | 46.8 |
| Terminal-Bench Hard | 31.1 | 1.5 |
| VideoMMMU | 82 | 69.4 |
| τ²-Bench Telecom (AA run) | 93.6 | 15.5 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.