Gemini 3.1 Pro vs GPT-5.5
Wins 2 of 8 areas
Images and charts · Following instructions
Wins 5 of 8 areas
Coding · Agents · Reasoning · Math · Long documents
GPT-5.5 is the stronger all-rounder.Gemini 3.1 Pro is cheaper and better at images and charts.
Scores updated · 74 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- MathCompetition and research-level math05GPT-5.55 of 5 tests
- AgentsCarrying out multi-step tasks on its own15GPT-5.55 of 6 tests
- ReasoningHard problems that need careful thinking13GPT-5.53 of 6 tests · 2 ties
- CodingWriting and fixing software13GPT-5.53 of 4 tests
- Long documentsFinding answers in very long texts01GPT-5.51 of 1 test
- Images and chartsUnderstanding pictures, charts and video30Gemini 3.1 Pro3 of 5 tests · 2 ties
- Following instructionsDoing exactly what it is asked20Gemini 3.1 Pro2 of 2 tests
- FactsGetting facts right instead of making them up11Even1 each
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemini 3.1 Pro costs 60% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where GPT-5.5 pulls ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+45.7points ahead
- Real work tasks from 44 professionsGDPVal+28points ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+25.6points ahead
Where Gemini 3.1 Pro pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+38.1points ahead
- Follows detailed instructions exactlyLiveBench · Instruction Following+8.4points ahead
- Math problems shown in pictures and chartsMathVista+6points ahead
Every test, side by side
All 74 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGPT-5.5
- LiveBench · Agentic CodingGPT-5.5 by 9.944.154+9.9
- LMArena · WebDevGPT-5.5 by 66 rating points14471513+66 rating
- LiveBench · CodingGPT-5.5 by 5.776.582.2+5.7
- SciCodeGemini 3.1 Pro by 2.958.755.8+2.9
AgentsGPT-5.5
- GDPValGPT-5.5 by 2814.742.7+28
- AA IT-Bench SREGPT-5.5 by 15.530.345.8+15.5
- MCP AtlasGPT-5.5 by 6.169.275.3+6.1
- AA ApexAgentsGPT-5.5 by 5.73237.7+5.7
- OSWorld-VerifiedGPT-5.5 by 2.576.278.7+2.5
- BrowseCompGemini 3.1 Pro by 1.585.984.4+1.5
ReasoningGPT-5.5
- CritPtGPT-5.5 by 9.417.727.1+9.4
- ARC-AGI-2GPT-5.5 by 7.977.185+7.9
- LiveBench · ReasoningGPT-5.5 by 5.78489.7+5.7
- Humanity's Last ExamGemini 3.1 Pro by 1.24745.8+1.2
- GPQA Diamondtie94.193.5tie
- ARC-AGI-3tie0.40.4tie
FactsEven
- AA-Omniscience · Non-hallucinationGemini 3.1 Pro by 38.149.111+38.1
- AA-Omniscience · AccuracyGPT-5.5 by 3.154.958+3.1
Images and chartsGemini 3.1 Pro
- MathVistaGemini 3.1 Pro by 690.284.2+6
- MMMU-ProGemini 3.1 Pro by 2.582.479.9+2.5
- OCRBenchv2Gemini 3.1 Pro by 1.762.861.1+1.7
- BLINKtie79.178.3tie
- LMArena · Visiontie12961292tie
MathGPT-5.5
- FrontierMath Tier 4GPT-5.5 by 45.726.872.5+45.7
- FrontierMath Tiers 1-3 (v2)GPT-5.5 by 25.659.685.3+25.6
- LiveBench · MathematicsGPT-5.5 by 4.99195.9+4.9
- HMMT Feb. 2026GPT-5.5 by 3.894.798.5+3.8
- AIME 2026GPT-5.5 by 1.798.3100+1.7
Following instructionsGemini 3.1 Pro
- LiveBench · Instruction FollowingGemini 3.1 Pro by 8.479.170.7+8.4
- IFBenchGemini 3.1 Pro by 1.277.175.9+1.2
Other results43 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- DeepSWEGPT-5.5 by 601070+60
- DeepSWE 1.1GPT-5.5 by 551267+55
- GDPval-AA v2GPT-5.5 by 547 rating points9621509+547 rating
- GDPval-AA v2 EloGPT-5.5 by 526 rating points9651491+526 rating
- GDPval-AA (Elo)GPT-5.5 by 452 rating points13171769+452 rating
- CyberGymGPT-5.5 by 4338.881.8+43
- Program BenchGPT-5.5 by 31.339.570.8+31.3
- AA Agentic IndexGPT-5.5 by 2710.337.3+27
- frontiermath_tier_4_v1GPT-5.5 by 18.716.735.4+18.7
- OfficeQA ProGemini 3.1 Pro by 18.472.554.1+18.4
- NL2RepoGPT-5.5 by 17.333.450.7+17.3
- AA-OmniscienceGemini 3.1 Pro by 11.431.920.5+11.4
- Finance Agent v2GPT-5.5 by 8.84351.8+8.8
- AA IntelligenceGPT-5.5 by 8.729.738.4+8.7
- MMSIBench (circular)GPT-5.5 by 8.627.436+8.6
- GDP (Surge AI)GPT-5.5 by 8.216.724.9+8.2
- PostTrainBenchGPT-5.5 by 6.821.628.4+6.8
- ERQAGemini 3.1 Pro by 6.370.864.5+6.3
- HLE CalibrationGemini 3.1 Pro by 6.250.444.2+6.2
- Artificial Analysis Coding IndexGPT-5.5 by 6.168.874.9+6.1
- Finance Agent v1.1GPT-5.5 by 5.659.765.3+5.6
- matharena_visual_math_overallGPT-5.5 by 5.589.494.9+5.5
- HiL-Bench (Tools-allowed)GPT-5.5 by 4.435.339.7+4.4
- OR-Bench (FRR)GPT-5.5 by 4.22.56.7+4.2
- DynaMathGPT-5.5 by 3.872.175.9+3.8
- RealWorldQAGemini 3.1 Pro by 3.285.482.2+3.2
- livebench_data_analysisGPT-5.5 by 3.178.581.6+3.1
- MathVerse (Vision-Only)Gemini 3.1 Pro by 3.187.784.6+3.1
- AIRS-BenchGPT-5.5 by 38386+3
- ARC-AGI-1Gemini 3.1 Pro by 39895+3
- Office QA Pro [Multimodal]Gemini 3.1 Pro by 372.569.5+3
- ForecastBench with searchGPT-5.5 by 2.461.964.3+2.4
- EmbSpatial-BenchGemini 3.1 Pro by 2.384.281.9+2.3
- Legal Agent BenchmarkGPT-5.5 by 2.102.1+2.1
- livebench_languageGPT-5.5 by 285.487.4+2
- HMMT Nov. 2025GPT-5.5 by 1.794.896.5+1.7
- BabyVisionGPT-5.5 by 1.554.455.9+1.5
- ChartQAProtie70.269.4tie
- HLE (with tools)tie51.452.2tie
- LiveBenchtie79.980.7tie
- DUDEtie82.181.7tie
- Finance Agenttie59.760tie
- Prophet Arenatie0.10.1tie
Questions people ask
Which is better, Gemini 3.1 Pro or GPT-5.5?
GPT-5.5 wins five of the eight areas we test: coding, agents, reasoning, math and long documents. Gemini 3.1 Pro wins images and charts and following instructions, and costs 60% less. They are level on facts.
Which is better for coding?
GPT-5.5. It wins 3 of the 4 coding tests both models report; Gemini 3.1 Pro wins 1.
Which is cheaper?
Gemini 3.1 Pro costs $2.00 per million input tokens and $12.00 per million output tokens; GPT-5.5 costs $5.00 and $30.00. That makes Gemini 3.1 Pro about 60% cheaper for the same work.
How do you compare the two?
We use the 74 benchmark tests both models have published scores on. The verdict counts the 31 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 43 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.