The two are evenly matched.DeepSeek-R1 is better at reasoning; MiniMax M1 40K at following instructions.
Scores updated · 22 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where MiniMax M1 40K pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+6.4points ahead
- Keeps track of context across a multi-turn chatMulti-Challenge+4points ahead
- Questions about very long textsLongBench v2+2.7points ahead
Where DeepSeek-R1 pulls ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+3.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+2.6points ahead
- Reasons across sets of long documentsAA-LCR+2points ahead
Every test, side by side
All 22 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- SWE-bench VerifiedMiniMax M1 40K by 6.449.255.6+6.4
- Terminal-Bench HardDeepSeek-R1 by 3.86.12.3+3.8
- SciCodetie38.337.8tie
ReasoningDeepSeek-R1
- GPQA DiamondDeepSeek-R1 by 2.670.868.2+2.6
- Humanity's Last Examtie8.57.8tie
Long documentsEven
- LongBench v2MiniMax M1 40K by 2.758.361+2.7
- AA-LCRDeepSeek-R1 by 257.755.7+2
Following instructionsMiniMax M1 40K
- Multi-ChallengeMiniMax M1 40K by 440.744.7+4
- IFBenchMiniMax M1 40K by 2.23941.2+2.2
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- OpenAI-MRCR (128k)MiniMax M1 40K by 40.335.876.1+40.3
- τ²-Bench Telecom (AA run)MiniMax M1 40K by 20.211.431.6+20.2
- SimpleQADeepSeek-R1 by 12.230.117.9+12.2
- Artificial Analysis Coding IndexDeepSeek-R1 by 10.524.614.1+10.5
- LiveCodeBench (24/8~25/5)MiniMax M1 40K by 6.455.962.3+6.4
- AIME 2025MiniMax M1 40K by 4.67074.6+4.6
- AIME 2024MiniMax M1 40K by 3.579.883.3+3.5
- MMLU-ProDeepSeek-R1 by 3.48480.6+3.4
- FullStackBenchDeepSeek-R1 by 2.570.167.6+2.5
- AA IntelligenceDeepSeek-R1 by 1.411.410+1.4
- ZebraLogicMiniMax M1 40K by 1.478.780.1+1.4
- MATH-500 (EM)DeepSeek-R1 by 1.397.396+1.3
- LiveCodeBenchDeepSeek-R1 by 1.263.562.3+1.2
Questions people ask
Which is better, DeepSeek-R1 or MiniMax M1 40K?
DeepSeek-R1 and MiniMax M1 40K each win one of the four areas where both have results. DeepSeek-R1 wins reasoning; MiniMax M1 40K wins following instructions. They are level on coding and long documents.
Which is better for coding?
Neither. They win 1 coding test each of the 3 both models report, and 1 is a tie.
How do you compare the two?
We use the 22 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.