VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Kimi K2.5 vs Kimi K2.6

MoonshotvsMoonshot69 shared benchmarks463 head-to-head
BenchmarkKimi K2.5Kimi K2.6
AA Agentic Index52.858.7
AA Intelligence3645.1
AA-LCR7376.7
AA-Omniscience-7.36.4
aime_202695.896.4
apexAgents11.528.5
arena_elo14501461
arena_text_factuality14451455
arena_vision12471264
Artificial Analysis Coding Index46.861.8
baby_vision36.568.5
baby_vision_with_python40.568.5
browsecomp74.986.3
BrowseComp (Agent Swarm)78.486.3
BrowseComp (w/ Ctx)74.983.2
charxiv_rq78.786.7
charxiv_rq_with_python78.786.7
coding_arena_elo14361509
critpt3.18
deepsearchqa_accuracy77.183
deepsearchqa_f18992.5
FORTRESS (Adversarial)54.165.6
FORTRESS (Benign)98.397.2
frontiermath_tier_44.214.6
gdpval38.341.4
GDPval-AA v210091190
Global-MMLU-Lite8488.4
GPQA Diamond87.991.1
HLE50.237.5
HLE (with tools)51.854
hmmt_feb_202687.194.7
IFBench70.276
imo_answer_bench81.886
livebench69.172.2
LiveCodeBench v68589.6
mathvision84.293.2
mathvision_with_python8593.2
MCP Atlas6468.1
mcpmark29.555.9
MMLU-Pro87.187.1
mmmu_pro_with_python77.780.1
MMMU-Pro78.580.1
OJBench (python)54.760.6
OmniScience Accuracy35.232.8
OmniScience Non-Hallucination5060.7
OSWorld-Verified63.373.1
scicode4953.5
simpleqa_verified36.938.7
strongreject99.599.8
SWE-bench Multilingual7376.7
SWE-bench Pro53.858.6
SWE-bench Verified76.880.2
SWEBench Pro (Public)50.758.6
Tau 3 Banking14.220.6
TauBench V3 - Banking14.223.3
Terminal Bench 2.1 (Best Harness)51.371.3
Terminal-Bench 2.050.866.7
Terminal-Bench 2.145.765.9
Terminal-Bench Hard34.843.9
toolathlon27.850
V* (w/ python)86.996.9
vectara_answer_rate92.299.7
vectara_avg_summary_length112116.7
vectara_factual_consistency85.889.2
vectara_hallucination_rate14.210.8
WideSearch7980.8
WideSearch (item-f1)72.780.8
τ²-Bench Telecom (AA run)95.995.9
τ³-Bench6620.6

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.