Kimi K2 Instruct vs Qwen3 VL 4B (Reasoning)
24 SHARED BENCHMARKSAcross 24 shared benchmarks, Kimi K2 Instruct scores higher on 20 and Qwen3 VL 4B (Reasoning) on 3, with 1 level. The widest gap is τ²-Bench Telecom (AA run), where Kimi K2 Instruct scores 73.4 against 15.5.
MOONSHOTVSALIBABA24 SHARED20–3 HEAD-TO-HEAD
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.
Compare Kimi K2 Instruct withEVERY TRACKED PAIRING
▶Anthropic12 PAIRINGS
▶OpenAI11 PAIRINGS
▶Google11 PAIRINGS
▶Alibaba10 PAIRINGS
▶DeepSeek5 PAIRINGS
▶Moonshot3 PAIRINGS
▶Meta2 PAIRINGS
▶Z.ai2 PAIRINGS
▶MiniMax1 PAIRING
▶NVIDIA1 PAIRING
Compare Qwen3 VL 4B (Reasoning) withEVERY TRACKED PAIRING
▶Anthropic12 PAIRINGS
▶OpenAI11 PAIRINGS
▶Google11 PAIRINGS
▶Alibaba10 PAIRINGS
▶DeepSeek5 PAIRINGS
▶Moonshot3 PAIRINGS
▶Meta2 PAIRINGS
▶Z.ai2 PAIRINGS
▶MiniMax1 PAIRING
▶NVIDIA1 PAIRING