DeepSeek-V4-Pro vs Qwen3.5 35B A3B
Wins 7 of 7 areas
Coding · Agents · Reasoning · Facts · Math · Long documents · Following instructions
Wins 0 of 7 areas
—
DeepSeek-V4-Pro is the stronger all-rounder.
Scores updated · 46 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software80DeepSeek-V4-Pro8 of 8 tests
- AgentsCarrying out multi-step tasks on its own50DeepSeek-V4-Pro5 of 5 tests
- ReasoningHard problems that need careful thinking30DeepSeek-V4-Pro3 of 3 tests
- MathCompetition and research-level math30DeepSeek-V4-Pro3 of 3 tests
- FactsGetting facts right instead of making them up21DeepSeek-V4-Pro2 of 3 tests
- Long documentsFinding answers in very long texts10DeepSeek-V4-Pro1 of 1 test
- Following instructionsDoing exactly what it is asked10DeepSeek-V4-Pro1 of 1 test
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V4-Pro costs 42% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V4-Pro pulls ahead
- Real work tasks from 44 professionsGDPVal+27.3points ahead
- Customer-service tasks in a simulated bankτ-Bench V3 · Banking+25.2points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+23.2points ahead
Where Qwen3.5 35B A3B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+8.7points ahead
Every test, side by side
All 46 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V4-Pro
- Terminal-Bench 2.1DeepSeek-V4-Pro by 23.26440.8+23.2
- LMArena · WebDevDeepSeek-V4-Pro by 213 rating points14631250+213 rating
- Terminal-Bench HardDeepSeek-V4-Pro by 19.746.226.5+19.7
- LiveCodeBench v6DeepSeek-V4-Pro by 17.992.574.6+17.9
- SWE-bench MultilingualDeepSeek-V4-Pro by 15.976.260.3+15.9
- SciCodeDeepSeek-V4-Pro by 13.150.837.7+13.1
- SWE-bench VerifiedDeepSeek-V4-Pro by 11.480.669.2+11.4
- SWE-bench ProDeepSeek-V4-Pro by 10.855.444.6+10.8
AgentsDeepSeek-V4-Pro
- GDPValDeepSeek-V4-Pro by 27.332.95.6+27.3
- τ-Bench V3 · BankingDeepSeek-V4-Pro by 25.230.14.9+25.2
- BrowseCompDeepSeek-V4-Pro by 22.483.461+22.4
- AA IT-Bench SREDeepSeek-V4-Pro by 16.838.321.5+16.8
- MCP AtlasDeepSeek-V4-Pro by 11.273.662.4+11.2
ReasoningDeepSeek-V4-Pro
- Humanity's Last ExamDeepSeek-V4-Pro by 16.537.521+16.5
- CritPtDeepSeek-V4-Pro by 1212.90.9+12
- GPQA DiamondDeepSeek-V4-Pro by 4.388.884.5+4.3
FactsDeepSeek-V4-Pro
- AA-Omniscience · AccuracyDeepSeek-V4-Pro by 22.94320.1+22.9
- AA-Omniscience · Non-hallucinationQwen3.5 35B A3B by 8.75.914.6+8.7
- Vectara HHEM hallucination ratelower is betterDeepSeek-V4-Pro by 1.98.610.5+1.9
MathDeepSeek-V4-Pro
- HMMT Feb. 2026DeepSeek-V4-Pro by 13.495.281.8+13.4
- IMOAnswerBenchDeepSeek-V4-Pro by 1389.876.8+13
- AIME 2026DeepSeek-V4-Pro by 3.496.793.3+3.4
Following instructionsDeepSeek-V4-Pro
- IFBenchDeepSeek-V4-Pro by 476.572.5+4
Other results22 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceDeepSeek-V4-Pro by 37.4-10.7-48.1+37.4
- Terminal-Bench 2.0DeepSeek-V4-Pro by 27.467.940.5+27.4
- Tool-DecathlonDeepSeek-V4-Pro by 24.152.828.7+24.1
- Artificial Analysis Coding IndexDeepSeek-V4-Pro by 22.459.437+22.4
- NL2RepoDeepSeek-V4-Pro by 1838.520.5+18
- AA Agentic IndexDeepSeek-V4-Pro by 15.927.711.8+15.9
- AgentWorldBench - SWEDeepSeek-V4-Pro by 11.859.447.6+11.8
- AA IntelligenceDeepSeek-V4-Pro by 11.130.419.3+11.1
- AgentWorldBench - OSDeepSeek-V4-Pro by 7.463.756.3+7.4
- τ²-Bench Telecom (AA run)DeepSeek-V4-Pro by 796.289.2+7
- vectara_avg_summary_lengthDeepSeek-V4-Pro by 58.9 rating points153.894.9+58.9 rating
- AgentWorldBench - MCPDeepSeek-V4-Pro by 5.463.357.9+5.4
- AgentWorldBench - OverallDeepSeek-V4-Pro by 5.35347.7+5.3
- AgentWorldBench - TerminalDeepSeek-V4-Pro by 5.251.346.1+5.2
- HMMT Nov. 2025DeepSeek-V4-Pro by 5.294.489.2+5.2
- GPQA (unspecified)DeepSeek-V4-Pro by 3.687.884.2+3.6
- AgentWorldBench - WebDeepSeek-V4-Pro by 3.250.347.1+3.2
- vectara_answer_rateQwen3.5 35B A3B by 2.697.299.8+2.6
- MMLU-ProDeepSeek-V4-Pro by 2.287.585.3+2.2
- AgentWorldBench - AndroidDeepSeek-V4-Pro by 255.253.2+2
- vectara_factual_consistencyDeepSeek-V4-Pro by 1.991.489.5+1.9
- AgentWorldBench - SearchDeepSeek-V4-Pro by 1.627.626+1.6
Questions people ask
Which is better, DeepSeek-V4-Pro or Qwen3.5 35B A3B?
DeepSeek-V4-Pro wins all seven areas where both have results: coding, agents, reasoning, facts, math, long documents and following instructions. Qwen3.5 35B A3B wins none.
Which is better for coding?
DeepSeek-V4-Pro. It wins 8 of the 8 coding tests both models report; Qwen3.5 35B A3B wins none.
Which is cheaper?
DeepSeek-V4-Pro costs $0.43 per million input tokens and $0.87 per million output tokens; Qwen3.5 35B A3B costs $0.25 and $2.00. That makes DeepSeek-V4-Pro about 42% cheaper for the same work.
How do you compare the two?
We use the 46 benchmark tests both models have published scores on. The verdict counts the 24 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 22 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.