DeepSeek-V4-Pro vs Gemini 3.5 Flash
Wins 1 of 7 areas
Long documents
Wins 6 of 7 areas
Coding · Agents · Reasoning · Facts · Math · Following instructions
Gemini 3.5 Flash is the stronger all-rounder.DeepSeek-V4-Pro is cheaper and better at long documents.
Scores updated · 42 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software15Gemini 3.5 Flash5 of 7 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own15Gemini 3.5 Flash5 of 6 tests
- ReasoningHard problems that need careful thinking03Gemini 3.5 Flash3 of 5 tests · 2 ties
- FactsGetting facts right instead of making them up03Gemini 3.5 Flash3 of 3 tests
- MathCompetition and research-level math12Gemini 3.5 Flash2 of 5 tests · 2 ties
- Following instructionsDoing exactly what it is asked01Gemini 3.5 Flash1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts10DeepSeek-V4-Pro1 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V4-Pro costs 88% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemini 3.5 Flash pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+31.9points ahead
- Common-sense trick questionsSimpleBench+25.8points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+24.4points ahead
Where DeepSeek-V4-Pro pulls ahead
- Complex command-line tasks across many fieldsTerminal-Bench 4.0+8points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+5.3points ahead
- Fresh competition math problemsLiveBench · Mathematics+2.5points ahead
Every test, side by side
All 42 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 3.5 Flash
- Terminal-Bench 2.1Gemini 3.5 Flash by 14.76478.7+14.7
- LiveBench · CodingGemini 3.5 Flash by 8.27078.2+8.2
- LiveBench · Agentic CodingGemini 3.5 Flash by 6.442.649+6.4
- Terminal-Bench HardDeepSeek-V4-Pro by 5.346.240.9+5.3
- LMArena · WebDevGemini 3.5 Flash by 35 rating points14641499+35 rating
- SciCodeGemini 3.5 Flash by 3.150.853.9+3.1
- SWE-bench Protie55.455.1tie
AgentsGemini 3.5 Flash
- AA ApexAgentsGemini 3.5 Flash by 22.824.347.1+22.8
- MCP AtlasGemini 3.5 Flash by 1073.683.6+10
- Terminal-Bench 4.0DeepSeek-V4-Pro by 814.66.6+8
- GDPValGemini 3.5 Flash by 2.332.935.2+2.3
- τ-Bench V3 · BankingGemini 3.5 Flash by 2.130.132.2+2.1
- AA IT-Bench SREGemini 3.5 Flash by 238.340.3+2
ReasoningGemini 3.5 Flash
- SimpleBenchGemini 3.5 Flash by 25.850.976.7+25.8
- Humanity's Last ExamGemini 3.5 Flash by 5.237.542.7+5.2
- GPQA DiamondGemini 3.5 Flash by 3.488.892.2+3.4
- LiveBench · Reasoningtie82.782tie
- CritPttie12.913.1tie
FactsGemini 3.5 Flash
- AA-Omniscience · Non-hallucinationGemini 3.5 Flash by 31.95.937.8+31.9
- SimpleQA VerifiedGemini 3.5 Flash by 2046.266.2+20
- AA-Omniscience · AccuracyGemini 3.5 Flash by 8.44351.4+8.4
MathGemini 3.5 Flash
- FrontierMath Tier 4Gemini 3.5 Flash by 24.42.426.8+24.4
- FrontierMath Tiers 1-3 (v2)Gemini 3.5 Flash by 17.545.362.8+17.5
- LiveBench · MathematicsDeepSeek-V4-Pro by 2.590.788.2+2.5
- AIME 2026tie94.695tie
- HMMT Feb. 2026tie95.295.5tie
Following instructionsGemini 3.5 Flash
- LiveBench · Instruction FollowingGemini 3.5 Flash by 13.262.475.6+13.2
- IFBenchtie76.576.3tie
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceGemini 3.5 Flash by 31.9-10.721.2+31.9
- DeepSWEGemini 3.5 Flash by 29837+29
- Artificial Analysis Coding IndexGemini 3.5 Flash by 10.759.470.1+10.7
- GDPval-AA (Elo)Gemini 3.5 Flash by 102 rating points15541656+102 rating
- livebench_data_analysisDeepSeek-V4-Pro by 9.674.564.9+9.6
- Terminal-Bench 2.0Gemini 3.5 Flash by 8.367.976.2+8.3
- livebench_languageGemini 3.5 Flash by 6.578.184.6+6.5
- ToolathlonGemini 3.5 Flash by 4.751.856.5+4.7
- GDPval-AA v2Gemini 3.5 Flash by 43 rating points13061349+43 rating
- AA IntelligenceGemini 3.5 Flash by 2.230.432.6+2.2
- LiveBenchGemini 3.5 Flash by 1.473.675+1.4
- τ²-Bench Telecom (AA run)tie96.295.3tie
- AA Agentic Indextie27.727.3tie
Questions people ask
Which is better, DeepSeek-V4-Pro or Gemini 3.5 Flash?
Gemini 3.5 Flash wins six of the seven areas where both have results: coding, agents, reasoning, facts, math and following instructions. DeepSeek-V4-Pro wins long documents, and costs 88% less.
Which is better for coding?
Gemini 3.5 Flash. It wins 5 of the 7 coding tests both models report; DeepSeek-V4-Pro wins 1, and 1 is a tie.
Which is cheaper?
DeepSeek-V4-Pro costs $0.43 per million input tokens and $0.87 per million output tokens; Gemini 3.5 Flash costs $1.50 and $9.00. That makes DeepSeek-V4-Pro about 88% cheaper for the same work.
How do you compare the two?
We use the 42 benchmark tests both models have published scores on. The verdict counts the 29 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.