Gemini 2.0 Flash Exp is the stronger all-rounder.
Scores updated · 40 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 2.0 Flash Exp pulls ahead
- Math problems shown in pictures and chartsMathVista+11.7points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+11points ahead
- Reads text in imagesOCRBench+4points ahead
Where GPT-4o pulls ahead
- College exam questions with charts, maps and diagramsMMMU+1.6points ahead
Every test, side by side
All 40 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningGemini 2.0 Flash Exp
- GPQA DiamondGemini 2.0 Flash Exp by 1163.652.6+11
- Humanity's Last ExamGemini 2.0 Flash Exp by 2.34.11.8+2.3
Images and chartsGemini 2.0 Flash Exp
Other results33 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- OlympiadBenchGemini 2.0 Flash Exp by 20.946.125.2+20.9
- SimpleQAGPT-4o by 11.626.638.2+11.6
- M-LongDoc (multimodal long-document benchmark)GPT-4o by 1031.441.4+10
- AI2DGPT-4o by 9.185.194.2+9.1
- Aider-PolyglotGPT-4o by 8.522.230.7+8.5
- IFEvalGemini 2.0 Flash Exp by 7.488.481+7.4
- Ruler 32kGemini 2.0 Flash Exp by 6.995.788.8+6.9
- Arena HardGPT-4o by 6.672.779.3+6.6
- Ruler 16kGemini 2.0 Flash Exp by 6.195.189+6.1
- naturalquestions_closedbookGPT-4o by 5.344.349.6+5.3
- RULER 64KGemini 2.0 Flash Exp by 5.393.788.4+5.3
- MATHGemini 2.0 Flash Exp by 4.890.185.3+4.8
- MTOB eng → kalam (ChrF) half bookGPT-4o by 4.849.554.3+4.8
- Chinese SimpleQA (C-SimpleQA)Gemini 2.0 Flash Exp by 4.663.358.7+4.6
- MEGA-Bench_macroGemini 2.0 Flash Exp by 4.553.949.4+4.5
- IFEval (avg)Gemini 2.0 Flash Exp by 4.388.484.1+4.3
- Ruler 8kGemini 2.0 Flash Exp by 3.99692.1+3.9
- GSM8KGemini 2.0 Flash Exp by 3.794.690.9+3.7
- ChartQAGemini 2.0 Flash Exp by 2.688.385.7+2.6
- MTOB eng → kalam (ChrF) no contextGemini 2.0 Flash Exp by 2.312.29.9+2.3
- OpenBookQAGPT-4o by 2.294.696.8+2.2
- MMLUGPT-4o by 2.171.773.8+2.1
- MMLU-ProGemini 2.0 Flash Exp by 1.776.474.7+1.7
- NarrativeQAGPT-4o by 1.278.379.5+1.2
- Ruler 4kGPT-4o by 19697+1
- AA Intelligencetie8.27.3tie
- MTOB kalam → eng (BLEURT) half booktie57.558.3tie
- HumanEvaltie89.690.2tie
- MTOB kalam → eng (BLEURT) no contexttie33.833.2tie
- MBPP+ (EvalPlus-augmented)tie75.976.2tie
- ChartQA_relaxedtie88.388.1tie
- DocVQAtie92.992.8tie
- DROP (F1)tie89.389.2tie
Questions people ask
Which is better, Gemini 2.0 Flash Exp or GPT-4o?
Gemini 2.0 Flash Exp wins two of the three areas where both have results: reasoning and images and charts. GPT-4o wins none. They are level on coding.
Which is better for coding?
Neither. All 1 coding tests both models report are ties.
How do you compare the two?
We use the 40 benchmark tests both models have published scores on. The verdict counts the 7 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 33 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.