VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V4-Flash vs Kimi K2.5

DeepSeekvsMoonshot44 shared benchmarks2815 head-to-head
BenchmarkDeepSeek-V4-FlashKimi K2.5
AA Agentic Index61.352.8
AA Intelligence51.836
AA-LCR74.373
AA-Omniscience-14.3-7.3
aime_202695.895.8
arc_agi_18465.3
ARC-AGI-24611.8
arena_elo14391450
arena_text_factuality14341445
Artificial Analysis Coding Index69.146.8
browsecomp73.274.9
BrowseComp (context management)73.274.9
coding_arena_elo14311436
critpt16.63.1
cybergym76.741.3
gdpval52.938.3
GDPval-AA v211891009
Global-MMLU-Lite88.484
GPQA Diamond90.887.9
HLE38.650.2
HLE (with tools)45.151.8
hmmt_feb_202694.887.1
IFBench79.270.2
imo_answer_bench88.481.8
MCP Atlas6964
MMLU-Pro86.487.1
nl2repo54.232
OmniScience Accuracy40.435.2
OmniScience Non-Hallucination15.650
scicode49.949
simplebench46.346.8
simpleqa_verified4536.9
strongreject97.499.5
SWE-bench Multilingual73.373
SWE-bench Pro52.653.8
SWE-bench Verified7976.8
Tau 3 Banking22.914.2
TauBench V3 - Banking39.414.2
Terminal-Bench 2.056.950.8
Terminal-Bench 2.182.745.7
Terminal-Bench Hard38.634.8
toolathlon70.327.8
τ²-Bench Telecom (AA run)95.695.9
τ³-Bench22.966

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.