VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Qwen2.5 VL 32B Instruct or Qwen3 VL Thinking (8B)?
Across 12 shared benchmarks, Qwen2.5 VL 32B Instruct scores higher on 2 and Qwen3 VL Thinking (8B) on 10. The widest gap is OSWorld-Verified, where Qwen3 VL Thinking (8B) scores 33.9 against 5.9.

Qwen2.5 VL 32B Instruct vs Qwen3 VL Thinking (8B)

Across 12 shared benchmarks, Qwen2.5 VL 32B Instruct scores higher on 2 and Qwen3 VL Thinking (8B) on 10. The widest gap is OSWorld-Verified, where Qwen3 VL Thinking (8B) scores 33.9 against 5.9.

AlibabavsAlibaba12 shared benchmarks210 head-to-head
BenchmarkQwen2.5 VL 32B InstructQwen3 VL Thinking (8B)
CC-OCR77.176.3
GPQA Diamond4669.9
LVBench4955.8
mathvision4062.7
MathVista74.781.4
MMLU-Pro68.877.3
MMMU7073.5
MMMU-Pro49.560.4
MMStar69.575.3
OSWorld-Verified5.933.9
screenspot_pro_no_tools39.446.6
Video-MME77.971.8

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.