Claude 3.5 Sonnet is the stronger all-rounder.
Scores updated · 34 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Reasoning rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude 3.5 Sonnet pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+1.6points ahead
Where MiniMax Text 01 pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 34 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results33 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- LongBench v2 medium (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 14.841.956.7+14.8
- MTOB eng → kalam (ChrF) no contextClaude 3.5 Sonnet by 14.220.26+14.2
- LongBench v2 easy (w/o CoT)MiniMax Text 01 by 1446.960.9+14
- LongBench v2 medium (w/o CoT)MiniMax Text 01 by 1438.652.6+14
- LongBench v2 short (w/o CoT)MiniMax Text 01 by 12.846.158.9+12.8
- Chinese SimpleQA (C-SimpleQA)MiniMax Text 01 by 1255.467.4+12
- LongBench v2 overall (w/o CoT)MiniMax Text 01 by 11.94152.9+11.9
- LongBench v2 easy (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 10.955.266.1+10.9
- LongBench v2 hard (w/o CoT)MiniMax Text 01 by 10.637.347.9+10.6
- Arena HardMiniMax Text 01 by 9.979.289.1+9.9
- LongBench v2 overall (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 9.846.756.5+9.8
- LongBench v2 hard (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 941.550.5+9
- LongBench v2 short (w/ CoT) (Pass@1 with chain-of-thought)MiniMax Text 01 by 7.853.961.7+7.8
- HumanEvalClaude 3.5 Sonnet by 6.893.786.9+6.8
- LongBench v2 long (w/o CoT)MiniMax Text 01 by 6.53743.5+6.5
- MATHMiniMax Text 01 by 6.371.177.4+6.3
- SimpleQAClaude 3.5 Sonnet by 4.728.423.7+4.7
- MTOB eng → kalam (ChrF) full bookClaude 3.5 Sonnet by 4.155.651.6+4.1
- MBPP+ (EvalPlus-augmented)Claude 3.5 Sonnet by 3.475.171.7+3.4
- IFEvalMiniMax Text 01 by 2.686.589.1+2.6
- MTOB kalam → eng (BLEURT) no contextMiniMax Text 01 by 2.331.433.6+2.3
- MMLU-ProClaude 3.5 Sonnet by 1.977.675.7+1.9
- MTOB eng → kalam (ChrF) half bookClaude 3.5 Sonnet by 1.953.651.7+1.9
- GSM8KClaude 3.5 Sonnet by 1.696.494.8+1.6
- DROP (F1)Claude 3.5 Sonnet by 188.887.8+1
- IFEval (avg)Claude 3.5 Sonnet by 190.189.1+1
- RULER 128Ktie93.894.7tie
- RULER 64Ktie95.294.3tie
- Ruler 16ktie95.795.3tie
- Ruler 32ktie9595.4tie
- MMLUtie88.788.5tie
- Ruler 4ktie96.596.3tie
- Ruler 8ktie9696.1tie
Questions people ask
Which is better, Claude 3.5 Sonnet or MiniMax Text 01?
Claude 3.5 Sonnet wins the one area where both have results: reasoning. MiniMax Text 01 wins none.
How do you compare the two?
We use the 34 benchmark tests both models have published scores on. The verdict counts the 1 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 33 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.