VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Claude Sonnet 4.5 vs Kimi K2 Instruct

63 SHARED BENCHMARKS

Across 63 shared benchmarks, Claude Sonnet 4.5 scores higher on 62 and Kimi K2 Instruct on 1. The widest gap is FinSearchComp-T3, where Claude Sonnet 4.5 scores 44 against 10.4. Kimi K2 Instruct is the cheaper of the two on tracked API pricing ($0.60 against $3.00 per million input tokens).

ANTHROPICVSMOONSHOT63 SHARED621 HEAD-TO-HEAD
AA-LCR68.353.7
browsecomp24.114.1
critpt1.10
Frames8558.1
gdpval40.418.2
HealthBench44.243.8
HLE19.87.9
HMMT 202574.638.8
IFBench57.342
mmlu_redux95.692.7
MMLU-Pro88.282.5
Seal-053.425.2
vectara_hallucination_rate↓ lower is better1217.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.

Compare Claude Sonnet 4.5 withEVERY TRACKED PAIRING

Anthropic11 PAIRINGS
OpenAI11 PAIRINGS
Google11 PAIRINGS
Alibaba11 PAIRINGS
DeepSeek5 PAIRINGS
Moonshot3 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA1 PAIRING

Compare Kimi K2 Instruct withEVERY TRACKED PAIRING

Anthropic11 PAIRINGS
OpenAI11 PAIRINGS
Google11 PAIRINGS
Alibaba11 PAIRINGS
DeepSeek5 PAIRINGS
Moonshot3 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA1 PAIRING