DeepSeek-V4-Pro vs Qwen3.6 Plus
Wins 6 of 7 areas
Coding · Agents · Reasoning · Facts · Math · Following instructions
Wins 1 of 7 areas
Long documents
DeepSeek-V4-Pro is the stronger all-rounder.Qwen3.6 Plus is better at long documents.
Scores updated · 51 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software72DeepSeek-V4-Pro7 of 10 tests · 1 tie
- MathCompetition and research-level math50DeepSeek-V4-Pro5 of 5 tests
- ReasoningHard problems that need careful thinking30DeepSeek-V4-Pro3 of 4 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own20DeepSeek-V4-Pro2 of 3 tests · 1 tie
- Following instructionsDoing exactly what it is asked20DeepSeek-V4-Pro2 of 2 tests
- FactsGetting facts right instead of making them up21DeepSeek-V4-Pro2 of 3 tests
- Long documentsFinding answers in very long texts01Qwen3.6 Plus1 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V4-Pro costs 63% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where DeepSeek-V4-Pro pulls ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+16.6points ahead
- Code for real scientific research problemsSciCode+10.1points ahead
- Research-level physics problemsCritPt+10points ahead
Where Qwen3.6 Plus pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+59.5points ahead
- Programming problems, refreshed regularlyLiveBench · Coding+8.2points ahead
- Reasons across sets of long documentsAA-LCR+3.6points ahead
Every test, side by side
All 51 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V4-Pro
- SciCodeDeepSeek-V4-Pro by 10.150.840.7+10.1
- LiveBench · CodingQwen3.6 Plus by 8.27078.2+8.2
- LiveCodeBench v6DeepSeek-V4-Pro by 5.492.587.1+5.4
- Terminal-Bench 2.1DeepSeek-V4-Pro by 2.66461.4+2.6
- SWE-bench MultilingualDeepSeek-V4-Pro by 2.476.273.8+2.4
- Terminal-Bench HardDeepSeek-V4-Pro by 2.346.243.9+2.3
- SWE-bench VerifiedDeepSeek-V4-Pro by 1.880.678.8+1.8
- LiveBench · Agentic CodingDeepSeek-V4-Pro by 1.242.641.4+1.2
- SWE-bench ProQwen3.6 Plus by 1.255.456.6+1.2
- LMArena · WebDevtie14641461tie
AgentsDeepSeek-V4-Pro
- τ-Bench V3 · BankingDeepSeek-V4-Pro by 9.330.120.8+9.3
- GDPValDeepSeek-V4-Pro by 8.232.924.7+8.2
- MCP Atlastie73.674.1tie
ReasoningDeepSeek-V4-Pro
- CritPtDeepSeek-V4-Pro by 1012.92.9+10
- Humanity's Last ExamDeepSeek-V4-Pro by 9.737.527.8+9.7
- LiveBench · ReasoningDeepSeek-V4-Pro by 6.982.775.8+6.9
- GPQA Diamondtie88.888.2tie
FactsDeepSeek-V4-Pro
- AA-Omniscience · Non-hallucinationQwen3.6 Plus by 59.55.965.4+59.5
- AA-Omniscience · AccuracyDeepSeek-V4-Pro by 16.64326.4+16.6
- SimpleQA VerifiedDeepSeek-V4-Pro by 2.146.244.1+2.1
MathDeepSeek-V4-Pro
- HMMT Feb. 2026DeepSeek-V4-Pro by 7.495.287.8+7.4
- FrontierMath Tiers 1-3 (v2)DeepSeek-V4-Pro by 745.338.3+7
- LiveBench · MathematicsDeepSeek-V4-Pro by 790.783.7+7
- IMOAnswerBenchDeepSeek-V4-Pro by 689.883.8+6
- AIME 2026DeepSeek-V4-Pro by 1.496.795.3+1.4
Following instructionsDeepSeek-V4-Pro
- LiveBench · Instruction FollowingDeepSeek-V4-Pro by 4.162.458.3+4.1
- IFBenchDeepSeek-V4-Pro by 1.376.575.2+1.3
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Tool-DecathlonDeepSeek-V4-Pro by 1352.839.8+13
- ToolathlonDeepSeek-V4-Pro by 1251.839.8+12
- AA-OmniscienceQwen3.6 Plus by 11.6-10.70.9+11.6
- HLE (with tools)DeepSeek-V4-Pro by 9.46050.6+9.4
- AgentWorldBench - MCPDeepSeek-V4-Pro by 863.355.3+8
- Terminal-Bench 2.0DeepSeek-V4-Pro by 6.367.961.6+6.3
- AgentWorldBench - SearchDeepSeek-V4-Pro by 5.727.621.9+5.7
- Artificial Analysis Coding IndexDeepSeek-V4-Pro by 4.959.454.5+4.9
- livebench_data_analysisDeepSeek-V4-Pro by 4.674.569.9+4.6
- AA IntelligenceDeepSeek-V4-Pro by 3.430.427+3.4
- AgentWorldBench - OSDeepSeek-V4-Pro by 3.463.760.3+3.4
- livebench_languageDeepSeek-V4-Pro by 3.178.175+3.1
- LiveBenchDeepSeek-V4-Pro by 2.773.670.8+2.7
- AgentWorldBench - AndroidQwen3.6 Plus by 2.555.257.6+2.5
- AgentWorldBench - OverallDeepSeek-V4-Pro by 2.25350.8+2.2
- τ²-Bench Telecom (AA run)Qwen3.6 Plus by 1.596.297.7+1.5
- AA Agentic IndexQwen3.6 Plus by 1.327.729+1.3
- MMLU-ProQwen3.6 Plus by 187.588.5+1
- AgentWorldBench - Terminaltie51.350.6tie
- NL2Repotie38.537.9tie
- AgentWorldBench - Webtie50.350.8tie
- AgentWorldBench - SWEtie59.459.1tie
- HMMT Nov. 2025tie94.494.6tie
Questions people ask
Which is better, DeepSeek-V4-Pro or Qwen3.6 Plus?
DeepSeek-V4-Pro wins six of the seven areas where both have results: coding, agents, reasoning, facts, math and following instructions. Qwen3.6 Plus wins long documents.
Which is better for coding?
DeepSeek-V4-Pro. It wins 7 of the 10 coding tests both models report; Qwen3.6 Plus wins 2, and 1 is a tie.
Which is cheaper?
DeepSeek-V4-Pro costs $0.43 per million input tokens and $0.87 per million output tokens; Qwen3.6 Plus costs $0.50 and $3.00. That makes DeepSeek-V4-Pro about 63% cheaper for the same work.
How do you compare the two?
We use the 51 benchmark tests both models have published scores on. The verdict counts the 28 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.