DeepSeek-V4-Flash vs MiMo V2 Pro
Wins 5 of 6 areas
Coding · Agents · Reasoning · Long documents · Following instructions
Wins 0 of 6 areas
—
DeepSeek-V4-Flash is the stronger all-rounder.
Scores updated · 20 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking30DeepSeek-V4-Flash3 of 3 tests
- CodingWriting and fixing software31DeepSeek-V4-Flash3 of 5 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own10DeepSeek-V4-Flash1 of 1 test
- Long documentsFinding answers in very long texts10DeepSeek-V4-Flash1 of 1 test
- Following instructionsDoing exactly what it is asked10DeepSeek-V4-Flash1 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where DeepSeek-V4-Flash pulls ahead
- Research-level physics problemsCritPt+16.3points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+13.8points ahead
- Reasons across sets of long documentsAA-LCR+11.4points ahead
Where MiMo V2 Pro pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+61.7points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+5.3points ahead
Every test, side by side
All 20 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingDeepSeek-V4-Flash
- SciCodeDeepSeek-V4-Flash by 7.850.342.5+7.8
- Terminal-Bench HardMiMo V2 Pro by 5.335.640.9+5.3
- SWE-bench MultilingualDeepSeek-V4-Flash by 1.673.371.7+1.6
- SWE-bench VerifiedDeepSeek-V4-Flash by 17978+1
- LMArena · WebDevtie14301433tie
ReasoningDeepSeek-V4-Flash
- CritPtDeepSeek-V4-Flash by 16.316.60.3+16.3
- Humanity's Last ExamDeepSeek-V4-Flash by 8.238.630.4+8.2
- GPQA DiamondDeepSeek-V4-Flash by 3.890.887+3.8
FactsEven
- AA-Omniscience · Non-hallucinationMiMo V2 Pro by 61.78.370+61.7
- AA-Omniscience · AccuracyDeepSeek-V4-Flash by 13.840.426.6+13.8
Long documentsDeepSeek-V4-Flash
- AA-LCRDeepSeek-V4-Flash by 11.479.768.3+11.4
Following instructionsDeepSeek-V4-Flash
- IFBenchDeepSeek-V4-Flash by 10.479.268.8+10.4
Other results7 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Artificial Analysis Coding IndexDeepSeek-V4-Flash by 27.769.141.4+27.7
- AA Agentic IndexMiMo V2 Pro by 21.141.762.8+21.1
- AA-OmniscienceMiMo V2 Pro by 18.9-14.34.6+18.9
- PinchBenchDeepSeek-V4-Flash by 10.391.381+10.3
- AA IntelligenceDeepSeek-V4-Flash by 5.734.328.6+5.7
- Terminal-Bench 2.0tie56.957.1tie
- τ²-Bench Telecom (AA run)tie9595tie
Questions people ask
Which is better, DeepSeek-V4-Flash or MiMo V2 Pro?
DeepSeek-V4-Flash wins five of the six areas where both have results: coding, agents, reasoning, long documents and following instructions. MiMo V2 Pro wins none. They are level on facts.
Which is better for coding?
DeepSeek-V4-Flash. It wins 3 of the 5 coding tests both models report; MiMo V2 Pro wins 1, and 1 is a tie.
How do you compare the two?
We use the 20 benchmark tests both models have published scores on. The verdict counts the 13 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 7 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.