Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Gemini 1.5 Pro vs GPT-4o

Google · released

Wins 1 of 3 areas

Reasoning

OpenAI · released

Wins 1 of 3 areas

Images and charts

The two are evenly matched.Gemini 1.5 Pro is better at reasoning; GPT-4o at images and charts.

Scores updated · 55 tests both models report · How we compare

Where each one wins

Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.

What it costs

Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.

Gemini 1.5 ProNo current price is tracked.
GPT-4o$5.00 to read · $15.00 to write$20Price from Artificial Analysis

The biggest differences

The three tests each model wins by the widest margin. Scores are out of 100.

Where Gemini 1.5 Pro pulls ahead

  • Math problems shown in pictures and chartsMathVista+6.7points aheadGemini 1.5 Pro68.1GPT-4o61.4
  • Graduate-level biology, physics and chemistry questionsGPQA Diamond+6.3points aheadGemini 1.5 Pro58.9GPT-4o52.6
  • Very hard expert questions across many subjectsHumanity's Last Exam+2.8points aheadGemini 1.5 Pro4.6GPT-4o1.8

Where GPT-4o pulls ahead

  • College exam questions with charts, maps and diagramsMMMU+6.3points aheadGPT-4o72.2Gemini 1.5 Pro65.9
  • Code for real scientific research problemsSciCode+3.8points aheadGPT-4o33.3Gemini 1.5 Pro29.5
  • Harder college exam questions with imagesMMMU-Pro+1.3points aheadGPT-4o56.3Gemini 1.5 Pro55

Every test, side by side

All 55 tests both models report. The winning score is in its model's colour; marks a score checked independently.

CodingEven

Full coding ranking

ReasoningGemini 1.5 Pro

Full reasoning ranking

Images and chartsGPT-4o

Full images and charts ranking

Other results47 tests, not counted

Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.

Questions people ask

Which is better, Gemini 1.5 Pro or GPT-4o?

Gemini 1.5 Pro and GPT-4o each win one of the three areas where both have results. Gemini 1.5 Pro wins reasoning; GPT-4o wins images and charts. They are level on coding.

Which is better for coding?

Neither. They win 1 coding test each of the 2 both models report.

How do you compare the two?

We use the 55 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 47 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.