Gemini 2.5 Pro vs o4-mini
Wins 4 of 8 areas
Reasoning · Facts · Images and charts · Long documents
Wins 2 of 8 areas
Agents · Math
Gemini 2.5 Pro wins more areas, narrowly.o4-mini is cheaper and better at agents.
Scores updated · 48 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking31Gemini 2.5 Pro3 of 4 tests
- FactsGetting facts right instead of making them up31Gemini 2.5 Pro3 of 4 tests
- Images and chartsUnderstanding pictures, charts and video21Gemini 2.5 Pro2 of 4 tests · 1 tie
- Long documentsFinding answers in very long texts10Gemini 2.5 Pro1 of 1 test
- AgentsCarrying out multi-step tasks on its own02o4-mini2 of 2 tests
- MathCompetition and research-level math02o4-mini2 of 2 tests
- CodingWriting and fixing software11Even1 each · 1 tie
- Following instructionsDoing exactly what it is asked11Even1 each
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
o4-mini costs 51% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemini 2.5 Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+37.2points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+14.3points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+11.3points ahead
Where o4-mini pulls ahead
- Finds hard-to-locate facts by browsing the webBrowseComp+41.6points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+20points ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+11.5points ahead
Every test, side by side
All 48 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- Terminal-Bench HardGemini 2.5 Pro by 11.326.515.2+11.3
- SWE-bench Verifiedo4-mini by 4.963.268.1+4.9
- SciCodetie46.346.5tie
Agentso4-mini
- BrowseCompo4-mini by 41.69.951.5+41.6
- GDPValo4-mini by 25.4025.4+25.4
ReasoningGemini 2.5 Pro
- GPQA DiamondGemini 2.5 Pro by 5.283.678.4+5.2
- CritPtGemini 2.5 Pro by 22.60.6+2
- Humanity's Last ExamGemini 2.5 Pro by 1.51816.5+1.5
- ARC-AGI-2o4-mini by 1.24.96.1+1.2
FactsGemini 2.5 Pro
- SimpleQA VerifiedGemini 2.5 Pro by 37.25618.8+37.2
- AA-Omniscience · AccuracyGemini 2.5 Pro by 14.33924.8+14.3
- Vectara HHEM hallucination ratelower is betterGemini 2.5 Pro by 11.6718.6+11.6
- AA-Omniscience · Non-hallucinationo4-mini by 10.49.119.5+10.4
Images and chartsGemini 2.5 Pro
- LMArena · VisionGemini 2.5 Pro by 69 rating points12631194+69 rating
- MMMU-ProGemini 2.5 Pro by 5.774.969.2+5.7
- MMMUo4-mini by 279.681.6+2
- MathVistatie83.984.3tie
Matho4-mini
- FrontierMath Tiers 1-3 (v2)o4-mini by 11.524.636.1+11.5
- FrontierMath Tier 4o4-mini by 4.904.9+4.9
Following instructionsEven
- IFBencho4-mini by 2048.768.7+20
- Multi-ChallengeGemini 2.5 Pro by 10.653.643+10.6
Other results26 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA Agentic Indexo4-mini by 32.63.536.1+32.6
- HarmBencho4-mini by 31.665.497+31.6
- ARC-AGI-1o4-mini by 21.73758.7+21.7
- Artificial Analysis Coding IndexGemini 2.5 Pro by 21.146.725.6+21.1
- AA-OmniscienceGemini 2.5 Pro by 19.4-16.4-35.7+19.4
- imo_2025Gemini 2.5 Pro by 17.331.614.3+17.3
- vectara_factual_consistencyGemini 2.5 Pro by 11.69381.4+11.6
- livecodebench_hardGemini 2.5 Pro by 10.559.448.9+10.5
- AIME 2025o4-mini by 9.78392.7+9.7
- Aider-PolyglotGemini 2.5 Pro by 7.676.568.9+7.6
- usamo_2025Gemini 2.5 Pro by 5.324.419.1+5.3
- AIR-Bench 2024o4-mini by 4.973.678.5+4.9
- TAU-bench (retail)o4-mini by 4.86771.8+4.8
- livecodebench_mediumGemini 2.5 Pro by 4.490.686.2+4.4
- LiveCodeBencho4-mini by 3.274.277.4+3.2
- simple_safety_testso4-mini by 397100+3
- bbqGemini 2.5 Pro by 2.496.494+2.4
- vectara_avg_summary_lengtho4-mini by 21.3 rating points106.4127.7+21.3 rating
- frontiermath_tier_4_v1o4-mini by 2.14.26.3+2.1
- AA Intelligenceo4-mini by 1.71516.7+1.7
- τ²-Bench Telecom (AA run)o4-mini by 1.554.155.6+1.5
- anthropic_red_teamGemini 2.5 Pro by 1.399.598.2+1.3
- XSTestGemini 2.5 Pro by 1.398.797.4+1.3
- TAU-bench (airline)tie5049.2tie
- vectara_answer_ratetie99.199.2tie
- livecodebench_easytie98.898.8tie
Questions people ask
Which is better, Gemini 2.5 Pro or o4-mini?
Gemini 2.5 Pro wins four of the eight areas we test: reasoning, facts, images and charts and long documents. o4-mini wins agents and math, and costs 51% less. They are level on coding and following instructions.
Which is better for coding?
Neither. They win 1 coding test each of the 3 both models report, and 1 is a tie.
Which is cheaper?
Gemini 2.5 Pro costs $1.25 per million input tokens and $10.00 per million output tokens; o4-mini costs $1.10 and $4.40. That makes o4-mini about 51% cheaper for the same work.
How do you compare the two?
We use the 48 benchmark tests both models have published scores on. The verdict counts the 22 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 26 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.