Gemini 2.5 Flash vs Magistral Medium 1
Wins 5 of 6 areas
Coding · Agents · Reasoning · Long documents · Following instructions
Wins 0 of 6 areas
—
Gemini 2.5 Flash is the stronger all-rounder.
Scores updated · 21 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking30Gemini 2.5 Flash3 of 4 tests · 1 tie
- CodingWriting and fixing software10Gemini 2.5 Flash1 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own10Gemini 2.5 Flash1 of 1 test
- Long documentsFinding answers in very long texts10Gemini 2.5 Flash1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemini 2.5 Flash1 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 2.5 Flash pulls ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+25.2points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+11.1points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+5.7points ahead
Where Magistral Medium 1 pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+16points ahead
Every test, side by side
All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 2.5 Flash
- Terminal-Bench HardGemini 2.5 Flash by 4.513.69.1+4.5
- SciCodetie29.129.7tie
ReasoningGemini 2.5 Flash
- GPQA DiamondGemini 2.5 Flash by 11.17967.9+11.1
- ARC-AGI-2Gemini 2.5 Flash by 2.52.50+2.5
- Humanity's Last ExamGemini 2.5 Flash by 2.312.19.8+2.3
- CritPttie1.10.3tie
FactsEven
- AA-Omniscience · Non-hallucinationMagistral Medium 1 by 1624.740.7+16
- AA-Omniscience · AccuracyGemini 2.5 Flash by 5.725.920.3+5.7
Following instructionsGemini 2.5 Flash
- IFBenchGemini 2.5 Flash by 25.250.325.1+25.2
Other results10 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ARC-AGI-1Gemini 2.5 Flash by 26.232.36.1+26.2
- LiveCodeBenchGemini 2.5 Flash by 25.976.250.3+25.9
- Aider-PolyglotGemini 2.5 Flash by 14.861.947.1+14.8
- AIME 2024Gemini 2.5 Flash by 11.682.370.7+11.6
- τ²-Bench Telecom (AA run)Gemini 2.5 Flash by 8.531.623.1+8.5
- AIME 2025Gemini 2.5 Flash by 7.17264.9+7.1
- Artificial Analysis Coding IndexGemini 2.5 Flash by 6.222.216+6.2
- AA IntelligenceGemini 2.5 Flash by 413.19.1+4
- AA Agentic IndexGemini 2.5 Flash by 3.118.815.7+3.1
- AA-OmniscienceMagistral Medium 1 by 2.9-29.8-26.9+2.9
Questions people ask
Which is better, Gemini 2.5 Flash or Magistral Medium 1?
Gemini 2.5 Flash wins five of the six areas where both have results: coding, agents, reasoning, long documents and following instructions. Magistral Medium 1 wins none. They are level on facts.
Which is better for coding?
Gemini 2.5 Flash. It wins 1 of the 2 coding tests both models report; Magistral Medium 1 wins none, and 1 is a tie.
How do you compare the two?
We use the 21 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 10 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.