DeepSeek-V4-Flash vs DeepSeek-V4-Pro
Wins 3 of 7 areas
Reasoning · Long documents · Following instructions
Wins 3 of 7 areas
Coding · Agents · Facts
The two are evenly matched.DeepSeek-V4-Flash is better at reasoning and following instructions; DeepSeek-V4-Pro at coding and agents.
Scores updated · 77 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking41DeepSeek-V4-Flash4 of 5 tests
- Following instructionsDoing exactly what it is asked20DeepSeek-V4-Flash2 of 2 tests
- Long documentsFinding answers in very long texts10DeepSeek-V4-Flash1 of 1 test
- CodingWriting and fixing software36DeepSeek-V4-Pro6 of 10 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own24DeepSeek-V4-Pro4 of 6 tests
- FactsGetting facts right instead of making them up12DeepSeek-V4-Pro2 of 3 tests
- MathCompetition and research-level math22Even2 each · 2 ties
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
DeepSeek-V4-Pro costs 26% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where DeepSeek-V4-Flash pulls ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+22points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+14.7points ahead
- Real work tasks from 44 professionsGDPVal+14points ahead
Where DeepSeek-V4-Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+12.1points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+10.6points ahead
- Finds hard-to-locate facts by browsing the webBrowseComp+10.2points ahead
Every test, side by side
All 77 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V4-Pro
- Terminal-Bench 2.1DeepSeek-V4-Flash by 14.778.764+14.7
- Terminal-Bench HardDeepSeek-V4-Pro by 10.635.646.2+10.6
- LiveBench · CodingDeepSeek-V4-Flash by 57570+5
- LiveBench · Agentic CodingDeepSeek-V4-Flash by 4.246.842.6+4.2
- LMArena · WebDevDeepSeek-V4-Pro by 33 rating points14301463+33 rating
- SWE-bench MultilingualDeepSeek-V4-Pro by 2.973.376.2+2.9
- SWE-bench ProDeepSeek-V4-Pro by 2.852.655.4+2.8
- LiveCodeBench v6DeepSeek-V4-Pro by 1.990.692.5+1.9
- SWE-bench VerifiedDeepSeek-V4-Pro by 1.67980.6+1.6
- SciCodetie50.350.8tie
AgentsDeepSeek-V4-Pro
- GDPValDeepSeek-V4-Flash by 1446.932.9+14
- BrowseCompDeepSeek-V4-Pro by 10.273.283.4+10.2
- τ-Bench V3 · BankingDeepSeek-V4-Flash by 9.339.430.1+9.3
- AA IT-Bench SREDeepSeek-V4-Pro by 6.831.538.3+6.8
- MCP AtlasDeepSeek-V4-Pro by 4.66973.6+4.6
- Terminal-Bench 4.0DeepSeek-V4-Pro by 2.512.114.6+2.5
ReasoningDeepSeek-V4-Flash
- SimpleBenchDeepSeek-V4-Pro by 4.646.350.9+4.6
- LiveBench · ReasoningDeepSeek-V4-Flash by 3.986.682.7+3.9
- CritPtDeepSeek-V4-Flash by 3.716.612.9+3.7
- GPQA DiamondDeepSeek-V4-Flash by 290.888.8+2
- Humanity's Last ExamDeepSeek-V4-Flash by 1.138.637.5+1.1
FactsDeepSeek-V4-Pro
- SimpleQA VerifiedDeepSeek-V4-Pro by 12.134.146.2+12.1
- AA-Omniscience · AccuracyDeepSeek-V4-Pro by 2.640.443+2.6
- AA-Omniscience · Non-hallucinationDeepSeek-V4-Flash by 2.48.35.9+2.4
MathEven
- FrontierMath Tier 4DeepSeek-V4-Flash by 2224.42.4+22
- FrontierMath Tiers 1-3 (v2)DeepSeek-V4-Flash by 12.257.545.3+12.2
- LiveBench · MathematicsDeepSeek-V4-Pro by 3.986.890.7+3.9
- IMOAnswerBenchDeepSeek-V4-Pro by 1.488.489.8+1.4
- AIME 2026tie95.896.7tie
- HMMT Feb. 2026tie94.895.2tie
Following instructionsDeepSeek-V4-Flash
- LiveBench · Instruction FollowingDeepSeek-V4-Flash by 3.165.562.4+3.1
- IFBenchDeepSeek-V4-Flash by 2.779.276.5+2.7
Other results44 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- DeepSWEDeepSeek-V4-Flash by 41.654.412.8+41.6
- Codeforces (Rating)DeepSeek-V4-Pro by 296 rating points30523348+296 rating
- DSBench-HardDeepSeek-V4-Flash by 28.559.631.1+28.5
- DSBench-FullStackDeepSeek-V4-Flash by 26.968.741.8+26.9
- ToolathlonDeepSeek-V4-Flash by 18.570.351.8+18.5
- GDPval-AA (Elo)DeepSeek-V4-Pro by 159 rating points13951554+159 rating
- NL2RepoDeepSeek-V4-Flash by 15.754.238.5+15.7
- HLE (with tools)DeepSeek-V4-Pro by 14.945.160+14.9
- Toolathlon VerifiedDeepSeek-V4-Flash by 14.470.355.9+14.4
- AA Agentic IndexDeepSeek-V4-Flash by 1441.727.7+14
- AutomationBench PublicDeepSeek-V4-Flash by 12.325.112.8+12.3
- GDPval-AA v2DeepSeek-V4-Pro by 117 rating points11891306+117 rating
- Terminal-Bench 2.0DeepSeek-V4-Pro by 1156.967.9+11
- Artificial Analysis Coding IndexDeepSeek-V4-Flash by 9.769.159.4+9.7
- CyberGymDeepSeek-V4-Pro by 6.676.783.3+6.6
- Apex (Pass@1)DeepSeek-V4-Flash by 5.63327.4+5.6
- livebench_data_analysisDeepSeek-V4-Flash by 4.879.374.5+4.8
- OpenAI-MRCR (1M)DeepSeek-V4-Pro by 4.678.783.3+4.6
- Apex-Shortlist (with tools)DeepSeek-V4-Pro by 4.58286.5+4.5
- IMOAnswerBench (with tools)DeepSeek-V4-Flash by 4.289.685.4+4.2
- CorpusQA 1M (ACC)DeepSeek-V4-Flash by 460.556.5+4
- AA IntelligenceDeepSeek-V4-Flash by 3.934.330.4+3.9
- AA-OmniscienceDeepSeek-V4-Pro by 3.6-14.3-10.7+3.6
- Apex-Shortlist (no tools)DeepSeek-V4-Pro by 3.482.485.8+3.4
- CritPt (no tools)DeepSeek-V4-Pro by 3.410.614+3.4
- HLE (wo / w tools)DeepSeek-V4-Pro by 3.145.148.2+3.1
- τ³-Bench BankingDeepSeek-V4-Pro by 3.122.926+3.1
- ProfBench (Search)DeepSeek-V4-Pro by 2.95759.9+2.9
- SWEBench Pro PublicDeepSeek-V4-Pro by 2.852.655.4+2.8
- PinchBenchDeepSeek-V4-Flash by 2.791.388.6+2.7
- SciCode (subtask)DeepSeek-V4-Pro by 2.348.250.5+2.3
- Vals.ai Financial Agent 1.1 - with web searchDeepSeek-V4-Pro by 2.260.162.3+2.2
- LiveCodeBenchDeepSeek-V4-Pro by 1.991.693.5+1.9
- MMLU-ProDeepSeek-V4-Pro by 1.386.287.5+1.3
- StrongREJECTDeepSeek-V4-Pro by 1.297.498.6+1.2
- τ²-Bench Telecom (AA run)DeepSeek-V4-Pro by 1.29596.2+1.2
- livebench_languageDeepSeek-V4-Flash by 1.179.278.1+1.1
- GPQA (unspecified)tie88.587.8tie
- Agents' Last Examtie25.225.7tie
- TauBench V3 - Averagetie73.773.2tie
- Vals.ai Financial Agent 1.1 - without web searchtie58.458.9tie
- Apex Shortlist (Pass@1)tie85.785.5tie
- TauBench V3 - Retailtie89.188.9tie
- TauBench V3 - Airlinetie80.880.8tie
Questions people ask
Which is better, DeepSeek-V4-Flash or DeepSeek-V4-Pro?
DeepSeek-V4-Flash and DeepSeek-V4-Pro each win three of the seven areas where both have results. DeepSeek-V4-Flash wins reasoning, long documents and following instructions; DeepSeek-V4-Pro wins coding, agents and facts. DeepSeek-V4-Pro costs 26% less. They are level on math.
Which is better for coding?
DeepSeek-V4-Pro. It wins 6 of the 10 coding tests both models report; DeepSeek-V4-Flash wins 3, and 1 is a tie.
Which is cheaper?
DeepSeek-V4-Flash costs $0.44 per million input tokens and $1.32 per million output tokens; DeepSeek-V4-Pro costs $0.43 and $0.87. That makes DeepSeek-V4-Pro about 26% cheaper for the same work.
How do you compare the two?
We use the 77 benchmark tests both models have published scores on. The verdict counts the 33 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 44 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.