DeepSeek-V4-Pro vs Kimi K2.7 Code
Wins 4 of 7 areas
Agents · Reasoning · Facts · Following instructions
Wins 3 of 7 areas
Coding · Math · Long documents
DeepSeek-V4-Pro wins more areas, narrowly.Kimi K2.7 Code is better at coding.
Scores updated · 33 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- AgentsCarrying out multi-step tasks on its own31DeepSeek-V4-Pro3 of 4 tests
- Following instructionsDoing exactly what it is asked20DeepSeek-V4-Pro2 of 2 tests
- ReasoningHard problems that need careful thinking21DeepSeek-V4-Pro2 of 5 tests · 2 ties
- FactsGetting facts right instead of making them up21DeepSeek-V4-Pro2 of 3 tests
- CodingWriting and fixing software23Kimi K2.7 Code3 of 6 tests · 1 tie
- MathCompetition and research-level math12Kimi K2.7 Code2 of 3 tests
- Long documentsFinding answers in very long texts01Kimi K2.7 Code1 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V4-Pro costs 74% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where DeepSeek-V4-Pro pulls ahead
- Complex command-line tasks across many fieldsTerminal-Bench 4.0+13.6points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+13.4points ahead
- Fresh competition math problemsLiveBench · Mathematics+11.1points ahead
Where Kimi K2.7 Code pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+11.7points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+9.8points ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+8.7points ahead
Every test, side by side
All 33 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingKimi K2.7 Code
- LiveBench · CodingKimi K2.7 Code by 47074+4
- Terminal-Bench 2.1Kimi K2.7 Code by 3.46467.4+3.4
- LiveBench · Agentic CodingKimi K2.7 Code by 3.142.645.7+3.1
- SciCodeDeepSeek-V4-Pro by 350.847.8+3
- Terminal-Bench HardDeepSeek-V4-Pro by 1.546.244.7+1.5
- LMArena · WebDevtie14641473tie
AgentsDeepSeek-V4-Pro
- Terminal-Bench 4.0DeepSeek-V4-Pro by 13.614.61+13.6
- τ-Bench V3 · BankingDeepSeek-V4-Pro by 9.930.120.2+9.9
- GDPValDeepSeek-V4-Pro by 5.932.927+5.9
- MCP AtlasKimi K2.7 Code by 2.473.676+2.4
ReasoningDeepSeek-V4-Pro
- SimpleBenchKimi K2.7 Code by 750.957.9+7
- CritPtDeepSeek-V4-Pro by 2.912.910+2.9
- Humanity's Last ExamDeepSeek-V4-Pro by 2.537.535+2.5
- GPQA Diamondtie88.889.6tie
- LiveBench · Reasoningtie82.782.8tie
FactsDeepSeek-V4-Pro
- AA-Omniscience · Non-hallucinationKimi K2.7 Code by 11.75.917.6+11.7
- SimpleQA VerifiedDeepSeek-V4-Pro by 9.746.236.5+9.7
- AA-Omniscience · AccuracyDeepSeek-V4-Pro by 3.44339.6+3.4
MathKimi K2.7 Code
- LiveBench · MathematicsDeepSeek-V4-Pro by 11.190.779.6+11.1
- FrontierMath Tier 4Kimi K2.7 Code by 9.82.412.2+9.8
- FrontierMath Tiers 1-3 (v2)Kimi K2.7 Code by 8.745.354+8.7
Following instructionsDeepSeek-V4-Pro
- IFBenchDeepSeek-V4-Pro by 13.476.563.1+13.4
- LiveBench · Instruction FollowingDeepSeek-V4-Pro by 6.162.456.3+6.1
Other results9 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- livebench_data_analysisDeepSeek-V4-Pro by 11.874.562.7+11.8
- τ²-Bench Telecom (AA run)DeepSeek-V4-Pro by 6.196.290.1+6.1
- Program BenchKimi K2.7 Code by 5.847.853.6+5.8
- AA Agentic IndexDeepSeek-V4-Pro by 5.227.722.5+5.2
- AA IntelligenceDeepSeek-V4-Pro by 4.630.425.8+4.6
- LiveBenchDeepSeek-V4-Pro by 1.773.671.9+1.7
- Artificial Analysis Coding IndexKimi K2.7 Code by 1.459.460.8+1.4
- AA-Omnisciencetie-10.7-10.2tie
- livebench_languagetie78.177.9tie
Questions people ask
Which is better, DeepSeek-V4-Pro or Kimi K2.7 Code?
DeepSeek-V4-Pro wins four of the seven areas where both have results: agents, reasoning, facts and following instructions. Kimi K2.7 Code wins coding, math and long documents.
Which is better for coding?
Kimi K2.7 Code. It wins 3 of the 6 coding tests both models report; DeepSeek-V4-Pro wins 2, and 1 is a tie.
Which is cheaper?
DeepSeek-V4-Pro costs $0.43 per million input tokens and $0.87 per million output tokens; Kimi K2.7 Code costs $0.95 and $4.00. That makes DeepSeek-V4-Pro about 74% cheaper for the same work.
How do you compare the two?
We use the 33 benchmark tests both models have published scores on. The verdict counts the 24 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 9 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.