Claude Opus 4.7 vs Gemini 3.1 Pro
Wins 3 of 8 areas
Coding · Agents · Math
Wins 4 of 8 areas
Reasoning · Facts · Images and charts · Following instructions
Gemini 3.1 Pro wins more areas, narrowly.Claude Opus 4.7 is better at coding.
Both rank among the ten best models we track in images and charts.
Scores updated · 88 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking15Gemini 3.1 Pro5 of 7 tests · 1 tie
- Images and chartsUnderstanding pictures, charts and video24Gemini 3.1 Pro4 of 6 tests
- FactsGetting facts right instead of making them up13Gemini 3.1 Pro3 of 4 tests
- Following instructionsDoing exactly what it is asked02Gemini 3.1 Pro2 of 2 tests
- CodingWriting and fixing software72Claude Opus 4.77 of 9 tests
- AgentsCarrying out multi-step tasks on its own61Claude Opus 4.76 of 7 tests
- MathCompetition and research-level math32Claude Opus 4.73 of 6 tests · 1 tie
- Long documentsFinding answers in very long texts11Even1 each
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemini 3.1 Pro costs 53% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemini 3.1 Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+21.8points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+18.5points ahead
- Common-sense trick questionsSimpleBench+17.9points ahead
Where Claude Opus 4.7 pulls ahead
- Finds one of eight look-alike replies in a long chatMRCR v2 8 needle 128k (average)+33points ahead
- Real work tasks from 44 professionsGDPVal+28.1points ahead
- Finds the root cause of IT system incidentsAA IT-Bench SRE+16.4points ahead
Every test, side by side
All 88 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Opus 4.7
- LMArena · WebDevClaude Opus 4.7 by 111 rating points15571446+111 rating
- SWE-bench ProClaude Opus 4.7 by 10.164.354.2+10.1
- Terminal-Bench 2.1Claude Opus 4.7 by 9.383.173.8+9.3
- SWE-bench VerifiedClaude Opus 4.7 by 787.680.6+7
- LiveBench · Agentic CodingClaude Opus 4.7 by 6.650.744.1+6.6
- LiveBench · CodingClaude Opus 4.7 by 5.682.176.5+5.6
- SciCodeGemini 3.1 Pro by 4.254.558.7+4.2
- SWE-bench MultilingualClaude Opus 4.7 by 3.680.576.9+3.6
- Terminal-Bench HardGemini 3.1 Pro by 2.351.553.8+2.3
AgentsClaude Opus 4.7
- GDPValClaude Opus 4.7 by 28.142.814.7+28.1
- AA IT-Bench SREClaude Opus 4.7 by 16.446.730.3+16.4
- τ-Bench V3 · BankingClaude Opus 4.7 by 13.234.621.4+13.2
- MCP AtlasClaude Opus 4.7 by 8.177.369.2+8.1
- BrowseCompGemini 3.1 Pro by 6.679.385.9+6.6
- AA ApexAgentsClaude Opus 4.7 by 1.933.932+1.9
- OSWorld-VerifiedClaude Opus 4.7 by 1.87876.2+1.8
ReasoningGemini 3.1 Pro
- SimpleBenchGemini 3.1 Pro by 17.961.779.6+17.9
- CritPtGemini 3.1 Pro by 5.71217.7+5.7
- Humanity's Last ExamGemini 3.1 Pro by 4.742.347+4.7
- LiveBench · ReasoningClaude Opus 4.7 by 3.287.284+3.2
- GPQA DiamondGemini 3.1 Pro by 2.791.494.1+2.7
- ARC-AGI-2Gemini 3.1 Pro by 1.375.877.1+1.3
- ARC-AGI-3tie0.20.4tie
FactsGemini 3.1 Pro
- SimpleQA VerifiedGemini 3.1 Pro by 21.851.773.5+21.8
- AA-Omniscience · Non-hallucinationClaude Opus 4.7 by 8.657.749.1+8.6
- AA-Omniscience · AccuracyGemini 3.1 Pro by 648.954.9+6
- Vectara HHEM hallucination ratelower is betterGemini 3.1 Pro by 1.61210.4+1.6
Images and chartsGemini 3.1 Pro
- BLINKGemini 3.1 Pro by 8.770.479.1+8.7
- CharXiv (RQ)Claude Opus 4.7 by 7.79183.3+7.7
- OCRBenchv2Gemini 3.1 Pro by 5.956.962.8+5.9
- MathVistaGemini 3.1 Pro by 5.884.490.2+5.8
- MMMU-ProGemini 3.1 Pro by 3.678.882.4+3.6
- LMArena · VisionClaude Opus 4.7 by 17 rating points13131296+17 rating
MathClaude Opus 4.7
- FrontierMath Tiers 1-3 (v2)Claude Opus 4.7 by 10.570.259.6+10.5
- USAMO 2026Gemini 3.1 Pro by 5.169.374.4+5.1
- FrontierMath Tier 4Claude Opus 4.7 by 4.931.726.8+4.9
- AIME 2026Gemini 3.1 Pro by 2.595.898.3+2.5
- LiveBench · MathematicsClaude Opus 4.7 by 1.992.891+1.9
- HMMT Feb. 2026tie93.994.7tie
Long documentsEven
- MRCR v2 8 needle 128k (average)Claude Opus 4.7 by 3359.326.3+33
- AA-LCRGemini 3.1 Pro by 3.378.782+3.3
Following instructionsGemini 3.1 Pro
- IFBenchGemini 3.1 Pro by 18.558.677.1+18.5
- LiveBench · Instruction FollowingGemini 3.1 Pro by 12.466.779.1+12.4
Other results45 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- DeepSWEClaude Opus 4.7 by 445410+44
- GDPval-AA (Elo)Claude Opus 4.7 by 436 rating points17531317+436 rating
- CyberGymClaude Opus 4.7 by 34.373.138.8+34.3
- BabyVisionGemini 3.1 Pro by 32.222.254.4+32.2
- AA Agentic IndexClaude Opus 4.7 by 29.239.510.3+29.2
- ERQAGemini 3.1 Pro by 18.352.570.8+18.3
- SimpleVQAGemini 3.1 Pro by 13.456.569.9+13.4
- VisuLogicGemini 3.1 Pro by 12.232.644.8+12.2
- AA IntelligenceClaude Opus 4.7 by 1140.729.7+11
- ToolathlonClaude Opus 4.7 by 10.559.348.8+10.5
- MathVerse (Vision-Only)Gemini 3.1 Pro by 10.377.487.7+10.3
- SWE-Pro BenchClaude Opus 4.7 by 10.164.354.2+10.1
- MMSIBench (circular)Gemini 3.1 Pro by 1017.427.4+10
- RealWorldQAGemini 3.1 Pro by 9.875.685.4+9.8
- DynaMathGemini 3.1 Pro by 963.172.1+9
- Finance Agent v2Claude Opus 4.7 by 8.551.543+8.5
- WorldVQAGemini 3.1 Pro by 8.435.944.3+8.4
- OfficeQA ProClaude Opus 4.7 by 8.180.672.5+8.1
- CharXiv Reasoning (With tools)Claude Opus 4.7 by 7.89183.2+7.8
- livebench_languageGemini 3.1 Pro by 7.577.985.4+7.5
- Legal Agent BenchmarkClaude Opus 4.7 by 7.17.10+7.1
- EmbSpatial-BenchGemini 3.1 Pro by 777.284.2+7
- τ²-Bench Telecom (AA run)Gemini 3.1 Pro by 788.695.6+7
- HiL-Bench (Tools-allowed)Claude Opus 4.7 by 6.441.735.3+6.4
- frontiermath_tier_4_v1Claude Opus 4.7 by 6.222.916.7+6.2
- ARC-AGI-1Gemini 3.1 Pro by 69298+6
- Artificial Analysis Coding IndexClaude Opus 4.7 by 4.873.668.8+4.8
- ZeroBench (sub)Gemini 3.1 Pro by 4.837.141.9+4.8
- ChartQAProGemini 3.1 Pro by 4.765.570.2+4.7
- Finance AgentClaude Opus 4.7 by 4.764.459.7+4.7
- Finance Agent v1.1Claude Opus 4.7 by 4.764.459.7+4.7
- AA-OmniscienceGemini 3.1 Pro by 4.627.331.9+4.6
- vectara_avg_summary_lengthClaude Opus 4.7 by 41.4 rating points149.1107.7+41.4 rating
- Office QA Pro [Multimodal]Claude Opus 4.7 by 476.572.5+4
- ZeroBench (main)Gemini 3.1 Pro by 4812+4
- HLE (with tools)Claude Opus 4.7 by 3.354.751.4+3.3
- LiveBenchGemini 3.1 Pro by 376.979.9+3
- Humanity’s Last ExamClaude Opus 4.7 by 2.546.944.4+2.5
- DUDEClaude Opus 4.7 by 2.484.582.1+2.4
- Blueprint-Bench 2Gemini 3.1 Pro by 224.526.5+2
- vectara_factual_consistencyGemini 3.1 Pro by 1.68889.6+1.6
- vectara_answer_rateGemini 3.1 Pro by 1.49899.4+1.4
- MMMLUGemini 3.1 Pro by 1.191.592.6+1.1
- Terminal-Bench 2.0tie69.468.5tie
- livebench_data_analysistie78.378.5tie
Questions people ask
Which is better, Claude Opus 4.7 or Gemini 3.1 Pro?
Gemini 3.1 Pro wins four of the eight areas we test: reasoning, facts, images and charts and following instructions. Claude Opus 4.7 wins coding, agents and math. They are level on long documents.
Which is better for coding?
Claude Opus 4.7. It wins 7 of the 9 coding tests both models report; Gemini 3.1 Pro wins 2.
Which is cheaper?
Claude Opus 4.7 costs $5.00 per million input tokens and $25.00 per million output tokens; Gemini 3.1 Pro costs $2.00 and $12.00. That makes Gemini 3.1 Pro about 53% cheaper for the same work.
How do you compare the two?
We use the 88 benchmark tests both models have published scores on. The verdict counts the 43 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 45 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.