DeepSeek-V4-Pro vs Kimi K2.6
Wins 4 of 7 areas
Agents · Reasoning · Facts · Math
Wins 3 of 7 areas
Coding · Long documents · Following instructions
DeepSeek-V4-Pro wins more areas, narrowly.Kimi K2.6 is better at coding.
Scores updated · 83 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- AgentsCarrying out multi-step tasks on its own52DeepSeek-V4-Pro5 of 7 tests
- MathCompetition and research-level math42DeepSeek-V4-Pro4 of 7 tests · 1 tie
- FactsGetting facts right instead of making them up31DeepSeek-V4-Pro3 of 4 tests
- ReasoningHard problems that need careful thinking21DeepSeek-V4-Pro2 of 4 tests · 1 tie
- CodingWriting and fixing software25Kimi K2.65 of 10 tests · 3 ties
- Following instructionsDoing exactly what it is asked01Kimi K2.61 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Kimi K2.61 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V4-Pro costs 74% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where DeepSeek-V4-Pro pulls ahead
- Complex command-line tasks across many fieldsTerminal-Bench 4.0+14.1points ahead
- Short factual questions, answered correctlySimpleQA Verified+11.3points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+10.4points ahead
Where Kimi K2.6 pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+53.6points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+23.2points ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+11.9points ahead
Every test, side by side
All 83 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingKimi K2.6
- LiveBench · CodingKimi K2.6 by 8.67078.6+8.6
- LMArena · WebDevKimi K2.6 by 45 rating points14641509+45 rating
- LiveBench · Agentic CodingKimi K2.6 by 4.342.646.9+4.3
- SWE-bench ProKimi K2.6 by 3.255.458.6+3.2
- LiveCodeBench v6DeepSeek-V4-Pro by 2.992.589.6+2.9
- Terminal-Bench HardDeepSeek-V4-Pro by 2.346.243.9+2.3
- Terminal-Bench 2.1Kimi K2.6 by 1.96465.9+1.9
- SciCodetie50.851.5tie
- SWE-bench Multilingualtie76.276.7tie
- SWE-bench Verifiedtie80.680.2tie
AgentsDeepSeek-V4-Pro
- Terminal-Bench 4.0DeepSeek-V4-Pro by 14.114.60.5+14.1
- AA IT-Bench SREDeepSeek-V4-Pro by 7.138.331.2+7.1
- τ-Bench V3 · BankingDeepSeek-V4-Pro by 6.830.123.3+6.8
- GDPValDeepSeek-V4-Pro by 5.932.927+5.9
- MCP AtlasDeepSeek-V4-Pro by 5.573.668.1+5.5
- AA ApexAgentsKimi K2.6 by 4.224.328.5+4.2
- BrowseCompKimi K2.6 by 2.983.486.3+2.9
ReasoningDeepSeek-V4-Pro
- CritPtDeepSeek-V4-Pro by 4.912.98+4.9
- LiveBench · ReasoningDeepSeek-V4-Pro by 3.382.779.4+3.3
- GPQA DiamondKimi K2.6 by 2.388.891.1+2.3
- Humanity's Last Examtie37.537.5tie
FactsDeepSeek-V4-Pro
- AA-Omniscience · Non-hallucinationKimi K2.6 by 53.65.959.5+53.6
- SimpleQA VerifiedDeepSeek-V4-Pro by 11.346.234.9+11.3
- AA-Omniscience · AccuracyDeepSeek-V4-Pro by 10.44332.6+10.4
- Vectara HHEM hallucination ratelower is betterDeepSeek-V4-Pro by 2.28.610.8+2.2
MathDeepSeek-V4-Pro
- FrontierMath Tier 4Kimi K2.6 by 23.22.425.6+23.2
- FrontierMath Tiers 1-3 (v2)Kimi K2.6 by 11.945.357.2+11.9
- USAMO 2026DeepSeek-V4-Pro by 9.560.751.2+9.5
- LiveBench · MathematicsDeepSeek-V4-Pro by 6.490.784.3+6.4
- IMOAnswerBenchDeepSeek-V4-Pro by 3.889.886+3.8
- HMMT Feb. 2026DeepSeek-V4-Pro by 2.595.292.7+2.5
- AIME 2026tie96.796.4tie
Following instructionsKimi K2.6
- LiveBench · Instruction FollowingKimi K2.6 by 262.464.4+2
- IFBenchtie76.576tie
Other results48 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceKimi K2.6 by 16-10.75.3+16
- Apex-Shortlist (with tools)DeepSeek-V4-Pro by 13.386.573.2+13.3
- GDPval-AA v2DeepSeek-V4-Pro by 116 rating points13061190+116 rating
- Apex Shortlist (Pass@1)DeepSeek-V4-Pro by 1085.575.5+10
- livebench_data_analysisDeepSeek-V4-Pro by 9.474.565.1+9.4
- Apex-Shortlist (no tools)DeepSeek-V4-Pro by 8.485.877.4+8.4
- IMOAnswerBench (with tools)Kimi K2.6 by 8.385.493.7+8.3
- GDPval-AA (Elo)DeepSeek-V4-Pro by 72 rating points15541482+72 rating
- MCPAtlas Public (Pass@1)DeepSeek-V4-Pro by 773.666.6+7
- HLE (with tools)DeepSeek-V4-Pro by 66054+6
- TauBench V3 - RetailDeepSeek-V4-Pro by 688.982.9+6
- AA Agentic IndexDeepSeek-V4-Pro by 5.627.722.1+5.6
- τ³-Bench BankingDeepSeek-V4-Pro by 5.42620.6+5.4
- TauBench V3 - AirlineKimi K2.6 by 580.885.8+5
- CritPt (no tools)DeepSeek-V4-Pro by 4.9149.1+4.9
- Vals.ai Financial Agent 1.1 - without web searchDeepSeek-V4-Pro by 4.958.954+4.9
- LiveCodeBenchDeepSeek-V4-Pro by 3.993.589.6+3.9
- ProfBench (Search)DeepSeek-V4-Pro by 3.959.956+3.9
- vectara_avg_summary_lengthDeepSeek-V4-Pro by 37.1 rating points153.8116.7+37.1 rating
- AgentWorldBench - AndroidKimi K2.6 by 3.755.258.9+3.7
- Vals.ai Financial Agent 1.1 - with web searchDeepSeek-V4-Pro by 3.562.358.8+3.5
- AA IntelligenceDeepSeek-V4-Pro by 3.430.427+3.4
- Apex (Pass@1)DeepSeek-V4-Pro by 3.427.424+3.4
- GPQA (unspecified)Kimi K2.6 by 3.287.891+3.2
- SWEBench Pro PublicKimi K2.6 by 3.255.458.6+3.2
- livebench_languageDeepSeek-V4-Pro by 378.175.1+3
- AgentWorldBench - OSDeepSeek-V4-Pro by 2.963.760.8+2.9
- vectara_answer_rateKimi K2.6 by 2.597.299.7+2.5
- Artificial Analysis Coding IndexKimi K2.6 by 2.459.461.8+2.4
- vectara_factual_consistencyDeepSeek-V4-Pro by 2.291.489.2+2.2
- AgentWorldBench - MCPKimi K2.6 by 1.963.365.2+1.9
- ToolathlonDeepSeek-V4-Pro by 1.851.850+1.8
- PinchBenchKimi K2.6 by 1.688.690.2+1.6
- SciCode (subtask)Kimi K2.6 by 1.550.552+1.5
- LiveBenchDeepSeek-V4-Pro by 1.473.672.2+1.4
- AgentWorldBench - TerminalKimi K2.6 by 1.251.352.5+1.2
- StrongREJECTKimi K2.6 by 1.298.699.8+1.2
- Terminal-Bench 2.0DeepSeek-V4-Pro by 1.267.966.7+1.2
- Global-MMLU-Litetie89.388.4tie
- TauBench V3 - Averagetie73.272.4tie
- AgentWorldBench - SWEtie59.458.8tie
- IOI 2025tie580.1585tie
- AgentWorldBench - Overalltie5353.4tie
- MMLU-Protie87.587.1tie
- τ²-Bench Telecom (AA run)tie96.295.9tie
- BrowseComp (w/ Ctx)tie83.483.2tie
- AgentWorldBench - Searchtie27.627.5tie
- AgentWorldBench - Webtie50.350.2tie
Questions people ask
Which is better, DeepSeek-V4-Pro or Kimi K2.6?
DeepSeek-V4-Pro wins four of the seven areas where both have results: agents, reasoning, facts and math. Kimi K2.6 wins coding, long documents and following instructions.
Which is better for coding?
Kimi K2.6. It wins 5 of the 10 coding tests both models report; DeepSeek-V4-Pro wins 2, and 3 are ties.
Which is cheaper?
DeepSeek-V4-Pro costs $0.43 per million input tokens and $0.87 per million output tokens; Kimi K2.6 costs $0.95 and $4.00. That makes DeepSeek-V4-Pro about 74% cheaper for the same work.
How do you compare the two?
We use the 83 benchmark tests both models have published scores on. The verdict counts the 35 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 48 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.