Across 15 shared benchmarks, Qwen2.5 Omni 7B scores higher on 1 and Qwen3.5 35B A3B on 14. The widest gap is mathvision, where Qwen3.5 35B A3B scores 83.9 against 25. Tracked API pricing per million tokens: Qwen2.5 Omni 7B $0.10 in / $6.76 out, Qwen3.5 35B A3B $0.25 in / $2.00 out.
| Benchmark | Qwen2.5 Omni 7B | Qwen3.5 35B A3B |
|---|---|---|
| AI2D | 83.2 | 92.6 |
| GPQA Diamond | 30.8 | 84.5 |
| GSM8K | 88.7 | 90.1 |
| humaneval | 78.7 | 66.5 |
| mathvision | 25 | 83.9 |
| MathVista | 67.9 | 86.2 |
| MBPP | 0.7 | 70.8 |
| mmlu_redux | 71 | 93.3 |
| MMLU-Pro | 47 | 85.3 |
| MMMU | 59.2 | 81.4 |
| MMMU-Pro | 36.6 | 75.1 |
| MMStar | 64 | 81.9 |
| MV-Bench | 70.3 | 74.8 |
| ocrbench_v2 | 57.8 | 65.3 |
| RealWorldQA | 70.3 | 84.1 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (alibaba-official, alibaba-official), otherwise the lowest tracked offer.