VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Qwen3.5 397B A17B vs Qwen3 VL 235B A22B Reasoning

31 SHARED BENCHMARKS

Across 31 shared benchmarks, Qwen3.5 397B A17B scores higher on 31 and Qwen3 VL 235B A22B Reasoning on 0. The widest gap is Terminal-Bench Hard, where Qwen3.5 397B A17B scores 40.9 against 11.4. Tracked API pricing per million tokens: Qwen3.5 397B A17B $0.60 in / $3.60 out, Qwen3 VL 235B A22B Reasoning $0.40 in / $4.00 out.

ALIBABAVSALIBABA31 SHARED310 HEAD-TO-HEAD
AA-LCR72.763
AA-Omniscience-30.8-46.5
arena_vision12651206
CC-OCR8281.5
charxiv_rq80.866.1
critpt1.70
ERQA67.552.5
GPQA Diamond89.377.2
HLE2913.6
HMMT 202592.777.4
IFBench78.856.5
ifeval92.688.2
include85.680
MLVU86.783.8
mmlu_prox84.780.6
mmlu_redux94.993.7
MMLU-Pro87.883.8
MMMU8578.7
MMMU-Pro7969.3
MMStar83.878.7
RealWorldQA83.981.3
scicode4239.9
SimpleVQA67.161.3
supergpqa70.464.3
VideoMMMU84.780

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare Qwen3.5 397B A17B withEVERY TRACKED PAIRING

Anthropic12 PAIRINGS
OpenAI11 PAIRINGS
Google10 PAIRINGS
Alibaba10 PAIRINGS
DeepSeek4 PAIRINGS
Moonshot4 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA2 PAIRINGS

Compare Qwen3 VL 235B A22B Reasoning withEVERY TRACKED PAIRING

Anthropic12 PAIRINGS
OpenAI11 PAIRINGS
Google10 PAIRINGS
Alibaba10 PAIRINGS
DeepSeek4 PAIRINGS
Moonshot4 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA2 PAIRINGS