VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, Qwen3.5 122B A10B or Qwen3 VL 4B (Reasoning)?
Across 40 shared benchmarks, Qwen3.5 122B A10B scores higher on 38 and Qwen3 VL 4B (Reasoning) on 2. The widest gap is Artificial Analysis Coding Index, where Qwen3.5 122B A10B scores 45.7 against 6.7.

Qwen3.5 122B A10B vs Qwen3 VL 4B (Reasoning)

Across 40 shared benchmarks, Qwen3.5 122B A10B scores higher on 38 and Qwen3 VL 4B (Reasoning) on 2. The widest gap is Artificial Analysis Coding Index, where Qwen3.5 122B A10B scores 45.7 against 6.7.

AlibabavsAlibaba40 shared benchmarks382 head-to-head
BenchmarkQwen3.5 122B A10BQwen3 VL 4B (Reasoning)
AA Agentic Index21.314.4
AA Intelligence32.87.7
AA-LCR70.323
AA-Omniscience-41.5-68.4
AI2D93.384.9
Artificial Analysis Coding Index45.76.7
CC-OCR81.873.8
critpt0.90
EmbSpatial-Bench0.880.7
ERQA6247.3
gdpval24.313.8
GPQA Diamond86.664.1
HallusionBench67.664.1
HLE47.54.6
HMMT 202590.353.1
IFBench76.136.6
ifeval93.482.6
include82.864.6
LiveCodeBench v678.951.3
LVBench74.453.5
mathvision86.260
MathVista87.479.5
mmlu_prox82.265
mmlu_redux9486
MMLU-Pro86.773.6
MMMU-Pro76.957
MMStar82.973.2
MV-Bench76.669.3
OCRBench92.181.6
OmniScience Accuracy24.412.2
OmniScience Non-Hallucination12.98.2
OSWorld-Verified5831.4
RealWorldQA85.173.2
RefSpatial-Bench0.745.3
scicode4217.1
screenspot_pro_no_tools70.449.2
supergpqa67.146.8
Terminal-Bench Hard31.11.5
VideoMMMU8269.4
τ²-Bench Telecom (AA run)93.615.5

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.