Gemini 1.5 Pro vs GPT-4o mini
Wins 3 of 3 areas
Coding · Reasoning · Images and charts
Wins 0 of 3 areas
—
Gemini 1.5 Pro is the stronger all-rounder.
Scores updated · 35 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 1.5 Pro pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+25.5points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+16.3points ahead
- Harder college exam questions with imagesMMMU-Pro+13.5points ahead
Where GPT-4o mini pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 35 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 1.5 Pro
- SWE-bench VerifiedGemini 1.5 Pro by 25.534.28.7+25.5
- SciCodeGemini 1.5 Pro by 6.629.522.9+6.6
ReasoningGemini 1.5 Pro
- GPQA DiamondGemini 1.5 Pro by 16.358.942.6+16.3
- Humanity's Last Examtie4.64.2tie
Images and chartsGemini 1.5 Pro
Other results27 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AI2DGemini 1.5 Pro by 19.294.475.2+19.2
- InterGPS (test)Gemini 1.5 Pro by 18.358.239.9+18.3
- MMLU-ProGemini 1.5 Pro by 14.175.861.7+14.1
- SimpleQAGemini 1.5 Pro by 13.523.49.9+13.5
- MMLUGemini 1.5 Pro by 12.779.566.8+12.7
- Artificial Analysis Coding IndexGemini 1.5 Pro by 12.223.611.4+12.2
- AIR-Bench 2024Gemini 1.5 Pro by 1167.356.3+11
- naturalquestions_closedbookGemini 1.5 Pro by 745.538.5+7
- GSM8KGemini 1.5 Pro by 6.590.884.3+6.5
- TextVQA-valGPT-4o mini by 6.464.570.9+6.4
- POPE (test)Gemini 1.5 Pro by 5.789.383.6+5.7
- HarmBenchGPT-4o mini by 579.984.9+5
- LiveCodeBenchGPT-4o mini by 530.535.5+5
- DROPGPT-4o mini by 4.874.979.7+4.8
- MMBench (dev en)Gemini 1.5 Pro by 4.187.983.8+4.1
- OpenBookQAGemini 1.5 Pro by 3.295.292+3.2
- HumanEvalGPT-4o mini by 3.184.187.2+3.1
- XSTestGemini 1.5 Pro by 2.898.896+2.8
- MATHGemini 1.5 Pro by 2.382.580.2+2.3
- MMMU (val) (Pass@1)Gemini 1.5 Pro by 254.152.1+2
- ScienceQA (img-test)Gemini 1.5 Pro by 28684+2
- anthropic_red_teamGemini 1.5 Pro by 1.699.998.3+1.6
- Video-MME OverallGemini 1.5 Pro by 1.462.661.2+1.4
- AA IntelligenceGemini 1.5 Pro by 1.27.96.7+1.2
- NarrativeQAGPT-4o mini by 1.275.676.8+1.2
- MGSMtie87.587tie
- simple_safety_teststie97.597.8tie
Questions people ask
Which is better, Gemini 1.5 Pro or GPT-4o mini?
Gemini 1.5 Pro wins all three areas where both have results: coding, reasoning and images and charts. GPT-4o mini wins none.
Which is better for coding?
Gemini 1.5 Pro. It wins 2 of the 2 coding tests both models report; GPT-4o mini wins none.
How do you compare the two?
We use the 35 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 27 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.