Gemini 3.1 Pro vs Gemini 3 Pro
Wins 7 of 8 areas
Coding · Agents · Reasoning · Facts · Images and charts · Math · Following instructions
Wins 0 of 8 areas
—
Gemini 3.1 Pro is the stronger all-rounder.
Both rank among the ten best models we track in images and charts.
Scores updated · 64 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software60Gemini 3.1 Pro6 of 7 tests · 1 tie
- ReasoningHard problems that need careful thinking50Gemini 3.1 Pro5 of 5 tests
- AgentsCarrying out multi-step tasks on its own31Gemini 3.1 Pro3 of 4 tests
- FactsGetting facts right instead of making them up20Gemini 3.1 Pro2 of 4 tests · 2 ties
- Following instructionsDoing exactly what it is asked20Gemini 3.1 Pro2 of 2 tests
- Images and chartsUnderstanding pictures, charts and video21Gemini 3.1 Pro2 of 6 tests · 3 ties
- MathCompetition and research-level math21Gemini 3.1 Pro2 of 3 tests
- Long documentsFinding answers in very long texts11Even1 each
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemini 3.1 Pro pulls ahead
- Abstract visual puzzles that people can solveARC-AGI-2+46points ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+40.6points ahead
- Finds hard-to-locate facts by browsing the webBrowseComp+26.7points ahead
Where Gemini 3 Pro pulls ahead
- Finds one of eight look-alike replies in a long chatMRCR v2 8 needle 128k (average)+50.7points ahead
- Real work tasks from 44 professionsGDPVal+19.5points ahead
- Olympiad math problems with short answersIMOAnswerBench+2.1points ahead
Every test, side by side
All 64 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 3.1 Pro
- Terminal-Bench HardGemini 3.1 Pro by 12.153.841.7+12.1
- SWE-bench MultilingualGemini 3.1 Pro by 11.976.965+11.9
- SWE-bench ProGemini 3.1 Pro by 10.954.243.3+10.9
- SWE-bench VerifiedGemini 3.1 Pro by 4.480.676.2+4.4
- LiveCodeBench v6Gemini 3.1 Pro by 4.391.787.4+4.3
- SciCodeGemini 3.1 Pro by 2.658.756.1+2.6
- LMArena · WebDevtie14471440tie
AgentsGemini 3.1 Pro
- BrowseCompGemini 3.1 Pro by 26.785.959.2+26.7
- GDPValGemini 3 Pro by 19.514.734.2+19.5
- MCP AtlasGemini 3.1 Pro by 15.169.254.1+15.1
- AA ApexAgentsGemini 3.1 Pro by 13.63218.4+13.6
ReasoningGemini 3.1 Pro
- ARC-AGI-2Gemini 3.1 Pro by 4677.131.1+46
- CritPtGemini 3.1 Pro by 8.617.79.1+8.6
- Humanity's Last ExamGemini 3.1 Pro by 7.34739.7+7.3
- GPQA DiamondGemini 3.1 Pro by 3.394.190.8+3.3
- SimpleBenchGemini 3.1 Pro by 3.279.676.4+3.2
FactsGemini 3.1 Pro
- AA-Omniscience · Non-hallucinationGemini 3.1 Pro by 40.649.18.5+40.6
- Vectara HHEM hallucination ratelower is betterGemini 3.1 Pro by 3.210.413.6+3.2
- AA-Omniscience · Accuracytie54.955.8tie
- SimpleQA Verifiedtie73.572.9tie
Images and chartsGemini 3.1 Pro
- MMMU-ProGemini 3.1 Pro by 2.282.480.2+2.2
- CharXiv (RQ)Gemini 3.1 Pro by 1.983.381.4+1.9
- Video-MMEGemini 3 Pro by 1.786.788.4+1.7
- LMArena · Visiontie12961305tie
- OCRBenchv2tie62.863.4tie
- MathVistatie90.289.8tie
MathGemini 3.1 Pro
- HMMT Feb. 2026Gemini 3.1 Pro by 8.394.786.4+8.3
- AIME 2026Gemini 3.1 Pro by 6.698.391.7+6.6
- IMOAnswerBenchGemini 3 Pro by 2.18183.1+2.1
Long documentsEven
- MRCR v2 8 needle 128k (average)Gemini 3 Pro by 50.726.377+50.7
- AA-LCRGemini 3.1 Pro by 68276+6
Following instructionsGemini 3.1 Pro
- IFBenchGemini 3.1 Pro by 6.777.170.4+6.7
- Multi-ChallengeGemini 3.1 Pro by 5.771.465.7+5.7
Other results31 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- LiveCodeBench Pro (Elo)Gemini 3.1 Pro by 448 rating points28872439+448 rating
- AA Agentic IndexGemini 3 Pro by 41.710.352+41.7
- browsecomp_with_context_managerGemini 3.1 Pro by 26.785.959.2+26.7
- ARC-AGI-1Gemini 3.1 Pro by 239875+23
- Artificial Analysis Coding IndexGemini 3.1 Pro by 22.368.846.5+22.3
- DeepSearchQA (F1)Gemini 3.1 Pro by 18.781.963.2+18.7
- AA-OmniscienceGemini 3.1 Pro by 16.631.915.3+16.6
- Terminal-Bench 2.0Gemini 3.1 Pro by 14.368.554.2+14.3
- t2-benchGemini 3.1 Pro by 13.999.385.4+13.9
- ToolathlonGemini 3.1 Pro by 12.448.836.4+12.4
- GDPval-AA (Elo)Gemini 3.1 Pro by 122 rating points13171195+122 rating
- τ²-Bench Telecom (AA run)Gemini 3.1 Pro by 8.595.687.1+8.5
- LVBenchGemini 3 Pro by 7.366.273.5+7.3
- LiveBenchGemini 3.1 Pro by 6.579.973.4+6.5
- HLE (with tools)Gemini 3.1 Pro by 5.651.445.8+5.6
- matharena_visual_math_overallGemini 3.1 Pro by 5.289.484.2+5.2
- MATH-VisionGemini 3.1 Pro by 3.789.886.1+3.7
- vectara_factual_consistencyGemini 3.1 Pro by 3.289.686.4+3.2
- WorldVQAGemini 3 Pro by 3.144.347.4+3.1
- frontiermath_tier_4_v1Gemini 3 Pro by 2.116.718.8+2.1
- AA IntelligenceGemini 3.1 Pro by 1.729.728+1.7
- HMMT Nov. 2025Gemini 3.1 Pro by 1.594.893.3+1.5
- LongVideoBenchGemini 3 Pro by 1.276.577.7+1.2
- MMLU-Protie9190.1tie
- MMMLUtie92.691.8tie
- vectara_avg_summary_lengthtie107.7101.9tie
- MotionBenchtie69.970.3tie
- LiveCodeBenchtie91.792tie
- SimpleVQAtie69.969.7tie
- mrcr_v2_8needle_1m_pointwisetie26.326.3tie
- vectara_answer_ratetie99.499.4tie
Questions people ask
Which is better, Gemini 3.1 Pro or Gemini 3 Pro?
Gemini 3.1 Pro wins seven of the eight areas we test: coding, agents, reasoning, facts, images and charts, math and following instructions. Gemini 3 Pro wins none. They are level on long documents.
Which is better for coding?
Gemini 3.1 Pro. It wins 6 of the 7 coding tests both models report; Gemini 3 Pro wins none, and 1 is a tie.
Which is cheaper?
Gemini 3.1 Pro costs $2.00 per million input tokens and $12.00 per million output tokens; Gemini 3 Pro costs $2.00 and $12.00. For a million tokens read plus a million written, they cost about the same.
How do you compare the two?
We use the 64 benchmark tests both models have published scores on. The verdict counts the 33 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 31 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.