VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

GPT-5 vs Kimi K2 Instruct

68 SHARED BENCHMARKS

Across 68 shared benchmarks, GPT-5 scores higher on 61 and Kimi K2 Instruct on 7. The widest gap is BrowseComp w/ tools, where GPT-5 scores 54.9 against 7.4. Kimi K2 Instruct is the cheaper of the two on tracked API pricing ($0.60 against $1.25 per million input tokens).

OPENAIVSMOONSHOT68 SHARED617 HEAD-TO-HEAD
BenchmarkGPT-5MARGINKimi K2 Instruct
AA-LCR76.353.7
browsecomp54.914.1
critpt5.70
Frames8658.1
gdpval32.518.2
GSM8K-1.897.3
harmbench7097.4
HealthBench67.243.8
HLE28.57.9
HMMT 202593.338.8
humaneval93.494.5
IFBench73.142
mmlu_redux95.392.7
MMLU-Pro87.582.5
Seal-051.425.2
Tau2 airline62.656.5
Tau2 retail81.170.6
vectara_hallucination_rate↓ lower is better14.717.9

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare GPT-5 withEVERY TRACKED PAIRING

Anthropic12 PAIRINGS
OpenAI10 PAIRINGS
Google11 PAIRINGS
Alibaba11 PAIRINGS
DeepSeek5 PAIRINGS
Moonshot3 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA1 PAIRING

Compare Kimi K2 Instruct withEVERY TRACKED PAIRING

Anthropic12 PAIRINGS
OpenAI10 PAIRINGS
Google11 PAIRINGS
Alibaba11 PAIRINGS
DeepSeek5 PAIRINGS
Moonshot3 PAIRINGS
Meta2 PAIRINGS
Z.ai2 PAIRINGS
MiniMax1 PAIRING
NVIDIA1 PAIRING