VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Kimi K2 Instruct vs Qwen3 VL 4B (Reasoning)

24 SHARED BENCHMARKS

Across 24 shared benchmarks, Kimi K2 Instruct scores higher on 20 and Qwen3 VL 4B (Reasoning) on 3, with 1 level. The widest gap is τ²-Bench Telecom (AA run), where Kimi K2 Instruct scores 73.4 against 15.5.

MOONSHOTVSALIBABA24 SHARED203 HEAD-TO-HEAD

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.

Compare Kimi K2 Instruct withEVERY TRACKED PAIRING

Anthropic12 PAIRINGS
OpenAI11 PAIRINGS
Google11 PAIRINGS
Alibaba10 PAIRINGS
DeepSeek5 PAIRINGS
Moonshot3 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA1 PAIRING

Compare Qwen3 VL 4B (Reasoning) withEVERY TRACKED PAIRING

Anthropic12 PAIRINGS
OpenAI11 PAIRINGS
Google11 PAIRINGS
Alibaba10 PAIRINGS
DeepSeek5 PAIRINGS
Moonshot3 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA1 PAIRING