Gemini 1.5 Pro is the stronger all-rounder.
Scores updated · 31 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemini 1.5 Pro pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+10.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+9.8points ahead
- Code for real scientific research problemsSciCode+2.8points ahead
Where Qwen2.5 Instruct 72B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 31 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 1.5 Pro
- SWE-bench VerifiedGemini 1.5 Pro by 10.434.223.8+10.4
- SciCodeGemini 1.5 Pro by 2.829.526.7+2.8
ReasoningGemini 1.5 Pro
- GPQA DiamondGemini 1.5 Pro by 9.858.949.1+9.8
- Humanity's Last ExamGemini 1.5 Pro by 14.63.6+1
Other results27 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- LiveCodeBenchQwen2.5 Instruct 72B by 2530.555.5+25
- SimpleQAGemini 1.5 Pro by 14.323.49.1+14.3
- Artificial Analysis Coding IndexGemini 1.5 Pro by 11.723.611.9+11.7
- MGSMGemini 1.5 Pro by 11.387.576.2+11.3
- Chinese SimpleQA (C-SimpleQA)Gemini 1.5 Pro by 1159.448.4+11
- naturalquestions_closedbookGemini 1.5 Pro by 9.745.535.9+9.7
- BBHGemini 1.5 Pro by 9.489.279.8+9.4
- DROP (F1)Gemini 1.5 Pro by 8.689.280.6+8.6
- HellaSwagGemini 1.5 Pro by 8.593.384.8+8.5
- AIR-Bench 2024Gemini 1.5 Pro by 8.367.359+8.3
- HarmBenchGemini 1.5 Pro by 7.179.972.8+7.1
- MATHQwen2.5 Instruct 72B by 5.982.588.4+5.9
- IFEvalGemini 1.5 Pro by 5.389.484.1+5.3
- GSM8KQwen2.5 Instruct 72B by 590.895.8+5
- MMLU-ProGemini 1.5 Pro by 4.775.871.1+4.7
- Arena HardGemini 1.5 Pro by 4.185.381.2+4.1
- HumanEvalQwen2.5 Instruct 72B by 2.584.186.6+2.5
- MMLUGemini 1.5 Pro by 2.579.577+2.5
- simple_safety_testsQwen2.5 Instruct 72B by 2.597.5100+2.5
- IFEval (avg)Gemini 1.5 Pro by 2.289.487.2+2.2
- DROPQwen2.5 Instruct 72B by 1.874.976.7+1.8
- MBPP+ (EvalPlus-augmented)Qwen2.5 Instruct 72B by 1.675.477+1.6
- NarrativeQAGemini 1.5 Pro by 1.175.674.5+1.1
- OpenBookQAQwen2.5 Instruct 72B by 195.296.2+1
- XSTesttie98.897.9tie
- anthropic_red_teamtie99.999.6tie
- AA Intelligencetie7.97.7tie
Questions people ask
Which is better, Gemini 1.5 Pro or Qwen2.5 Instruct 72B?
Gemini 1.5 Pro wins all two areas where both have results: coding and reasoning. Qwen2.5 Instruct 72B wins none.
Which is better for coding?
Gemini 1.5 Pro. It wins 2 of the 2 coding tests both models report; Qwen2.5 Instruct 72B wins none.
How do you compare the two?
We use the 31 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 27 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.