Gemini 2.0 Flash (Reasoning) vs GPT-4o
Wins 2 of 3 areas
Reasoning · Images and charts
Wins 1 of 3 areas
Coding
Gemini 2.0 Flash (Reasoning) is the stronger all-rounder.GPT-4o is better at coding.
Scores updated · 13 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding, reasoning and images and charts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 2.0 Flash (Reasoning) pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+21.6points ahead
- College exam questions with charts, maps and diagramsMMMU+3.2points ahead
Where GPT-4o pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+19.7points ahead
Every test, side by side
All 13 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningGemini 2.0 Flash (Reasoning)
- GPQA DiamondGemini 2.0 Flash (Reasoning) by 21.674.252.6+21.6
Images and chartsGemini 2.0 Flash (Reasoning)
- MMMUGemini 2.0 Flash (Reasoning) by 3.275.472.2+3.2
Other results10 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HarmBenchGPT-4o by 16.766.282.9+16.7
- AIME 2025Gemini 2.0 Flash (Reasoning) by 15.827.511.7+15.8
- AIR-Bench 2024Gemini 2.0 Flash (Reasoning) by 3.866.262.4+3.8
- LiveCodeBenchGPT-4o by 3.235.138.3+3.2
- XSTestGPT-4o by 295.397.3+2
- MMLU-ProGemini 2.0 Flash (Reasoning) by 1.776.474.7+1.7
- EgoSchematie71.572.2tie
- anthropic_red_teamtie99.499.1tie
- bbqtie95.495.1tie
- simple_safety_teststie98.598.5tie
Questions people ask
Which is better, Gemini 2.0 Flash (Reasoning) or GPT-4o?
Gemini 2.0 Flash (Reasoning) wins two of the three areas where both have results: reasoning and images and charts. GPT-4o wins coding.
Which is better for coding?
GPT-4o. It wins the one coding test both models report.
How do you compare the two?
We use the 13 benchmark tests both models have published scores on. The verdict counts the 3 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 10 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.