MiniMax Text 01 is the stronger all-rounder.
Scores updated · 32 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where MiniMax Text 01 pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+1.8points ahead
Where GPT-4o pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 32 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results31 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MMLUMiniMax Text 01 by 14.773.888.5+14.7
- SimpleQAGPT-4o by 14.538.223.7+14.5
- LongBench v2 easy (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 11.954.266.1+11.9
- Arena HardMiniMax Text 01 by 9.879.389.1+9.8
- Chinese SimpleQA (C-SimpleQA)MiniMax Text 01 by 8.758.767.4+8.7
- IFEvalMiniMax Text 01 by 8.18189.1+8.1
- LongBench v2 medium (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 8.148.656.7+8.1
- MATHGPT-4o by 7.985.377.4+7.9
- Ruler 32kMiniMax Text 01 by 6.688.895.4+6.6
- Ruler 16kMiniMax Text 01 by 6.38995.3+6.3
- RULER 64KMiniMax Text 01 by 5.988.494.3+5.9
- LongBench v2 short (w/o CoT)MiniMax Text 01 by 5.653.358.9+5.6
- LongBench v2 overall (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 5.151.456.5+5.1
- IFEval (avg)MiniMax Text 01 by 584.189.1+5
- MBPP+ (EvalPlus-augmented)GPT-4o by 4.576.271.7+4.5
- Ruler 8kMiniMax Text 01 by 492.196.1+4
- GSM8KMiniMax Text 01 by 3.990.994.8+3.9
- MTOB eng → kalam (ChrF) no contextGPT-4o by 3.99.96+3.9
- LongBench v2 easy (w/o CoT)MiniMax Text 01 by 3.557.460.9+3.5
- HumanEvalGPT-4o by 3.390.286.9+3.3
- LongBench v2 long (w/o CoT)MiniMax Text 01 by 3.340.243.5+3.3
- LongBench v2 overall (w/o CoT)MiniMax Text 01 by 2.850.152.9+2.8
- MTOB eng → kalam (ChrF) half bookGPT-4o by 2.654.351.7+2.6
- LongBench v2 hard (w/o CoT)MiniMax Text 01 by 2.345.647.9+2.3
- LongBench v2 short (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 2.159.661.7+2.1
- DROP (F1)GPT-4o by 1.489.287.8+1.4
- MMLU-ProMiniMax Text 01 by 174.775.7+1
- LongBench v2 hard (w/ CoT) (Pass@1 with chain-of-thought)tie49.750.5tie
- Ruler 4ktie9796.3tie
- MTOB kalam → eng (BLEURT) no contexttie33.233.6tie
- LongBench v2 medium (w/o CoT)tie52.452.6tie
Questions people ask
Which is better, GPT-4o or MiniMax Text 01?
MiniMax Text 01 wins the one area where both have results: reasoning. GPT-4o wins none.
How do you compare the two?
We use the 32 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 31 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.