Across 12 shared benchmarks, Qwen2.5 Omni 7B scores higher on 1 and Qwen3 VL Thinking (8B) on 11. The widest gap is mathvision, where Qwen3 VL Thinking (8B) scores 62.7 against 25. Tracked API pricing per million tokens: Qwen2.5 Omni 7B $0.10 in / $6.76 out, Qwen3 VL Thinking (8B) $0.18 in / $2.10 out.
| Benchmark | Qwen2.5 Omni 7B | Qwen3 VL Thinking (8B) |
|---|---|---|
| AI2D | 83.2 | 84.9 |
| GPQA Diamond | 30.8 | 69.9 |
| mathvision | 25 | 62.7 |
| MathVista | 67.9 | 81.4 |
| MM MTBench | 0.1 | 8 |
| mmlu_redux | 71 | 88.8 |
| MMLU-Pro | 47 | 77.3 |
| MMMU | 59.2 | 73.5 |
| MMMU-Pro | 36.6 | 60.4 |
| MMStar | 64 | 75.3 |
| MV-Bench | 70.3 | 69 |
| RealWorldQA | 70.3 | 73.5 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (alibaba-official, alibaba-official), otherwise the lowest tracked offer.