VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek-V3.2 vs Kimi K2 (Reasoning)

DeepSeekvsMoonshot71 shared benchmarks2942 head-to-head
BenchmarkDeepSeek-V3.2Kimi K2 (Reasoning)
AA Agentic Index39.847.9
AA Intelligence32.833.5
AA-LCR70.770.3
AA-Omniscience-22.5-21.4
AIME 202593.194.5
AIME25 no tools89.394.5
AIME25 w/ python58.199.1
Arena-Hard (Creative Writing)88.880.1
Arena-Hard (Hard Prompt)53.471.9
ArtifactsBench55.854.2
Artificial Analysis Coding Index44.234.8
browsecomp51.460.2
BrowseComp w/ tools40.160.2
browsecomp_with_context_manager67.660.2
browsecomp_zh6562.3
BrowseComp-ZH w/ tools47.962.3
critpt2.92.6
FinSearchComp-global26.229.5
FinSearchComp-T32747.4
FinSearchComp-T3 w/ tools2747.4
Frames80.287
Frames w/ tools80.287
frontiermath_tier_42.10
GAIA (text only)63.560.2
gaia_no_file75.175.6
gdpval18.824.5
GPQA (unspecified)79.974.2
GPQA Diamond8484.5
HealthBench46.958
HealthBench no tools46.958
HLE40.823.9
HLE (with tools)40.844.9
HMMT 202590.289.4
hmmt_feb_202592.593.3
hmmt_nov_202590.289.2
HMMT25 no tools83.689.4
HMMT25 w/ python49.595.1
IFBench6168.1
imo_answer_bench78.378.6
LiveCodeBench8679.2
LiveCodeBench v683.383.1
LiveCodeBenchV6 no tools74.183.1
longbench_v259.845.1
Longform Writing eval (Kimi K2 Thinking system card)72.573.8
mmlu_redux93.794.4
MMLU-Pro8684.6
mrcr55.544.2
Multi-SWE-Bench37.441.9
OJ-Bench (cpp)54.748.7
OJ-Bench (cpp) no tools38.248.7
OmniScience Accuracy3330.9
OmniScience Non-Hallucination17.325.8
ResearchRubrics55.856.2
scicode3944.8
Seal-049.556.3
Seal-0 w/ tools38.556.3
swe_bench_bash6063.4
SWE-bench Multilingual70.261.1
SWE-bench Verified73.171.3
Terminal-Bench37.747.1
Terminal-Bench 2.046.435.7
Terminal-Bench Hard35.631.1
Terminal-Bench w/ simulated tools (JSON)37.747.1
theagentcompany3430
vectara_answer_rate92.698.6
vectara_avg_summary_length6259.2
vectara_factual_consistency93.782.1
vectara_hallucination_rate6.317.9
xbench_deepsearch7176
τ²-Bench9174.3
τ²-Bench Telecom (AA run)90.693

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.