Claude Opus 4.8 vs Gemini 3.1 Pro
Wins 5 of 8 areas
Coding · Agents · Reasoning · Images and charts · Math
Wins 3 of 8 areas
Facts · Long documents · Following instructions
Claude Opus 4.8 is the stronger all-rounder.Gemini 3.1 Pro is cheaper and better at following instructions.
Scores updated · 84 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software81Claude Opus 4.88 of 9 tests
- MathCompetition and research-level math60Claude Opus 4.86 of 7 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own61Claude Opus 4.86 of 7 tests
- ReasoningHard problems that need careful thinking43Claude Opus 4.84 of 7 tests
- Images and chartsUnderstanding pictures, charts and video10Claude Opus 4.81 of 2 tests · 1 tie
- Following instructionsDoing exactly what it is asked02Gemini 3.1 Pro2 of 2 tests
- FactsGetting facts right instead of making them up12Gemini 3.1 Pro2 of 3 tests
- Long documentsFinding answers in very long texts01Gemini 3.1 Pro1 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemini 3.1 Pro costs 53% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Claude Opus 4.8 pulls ahead
- Real work tasks from 44 professionsGDPVal+33.1points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+29.3points ahead
- Proof problems from the US Math OlympiadUSAMO 2026+22.3points ahead
Where Gemini 3.1 Pro pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+20.5points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+14.9points ahead
- Common-sense trick questionsSimpleBench+14.8points ahead
Every test, side by side
All 84 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Opus 4.8
- SWE-bench ProClaude Opus 4.8 by 1569.254.2+15
- LMArena · WebDevClaude Opus 4.8 by 109 rating points15561447+109 rating
- Terminal-Bench 2.1Claude Opus 4.8 by 10.884.673.8+10.8
- SWE-bench VerifiedClaude Opus 4.8 by 888.680.6+8
- SWE-bench MultilingualClaude Opus 4.8 by 7.584.476.9+7.5
- LiveBench · Agentic CodingClaude Opus 4.8 by 6.450.544.1+6.4
- LiveBench · CodingClaude Opus 4.8 by 5.381.876.5+5.3
- Terminal-Bench HardClaude Opus 4.8 by 4.558.353.8+4.5
- SciCodeGemini 3.1 Pro by 4.354.458.7+4.3
AgentsClaude Opus 4.8
- GDPValClaude Opus 4.8 by 33.147.814.7+33.1
- Terminal-Bench 4.0Claude Opus 4.8 by 17.721.74+17.7
- MCP AtlasClaude Opus 4.8 by 1382.269.2+13
- τ-Bench V3 · BankingClaude Opus 4.8 by 12.834.221.4+12.8
- AA ApexAgentsClaude Opus 4.8 by 7.439.432+7.4
- OSWorld-VerifiedClaude Opus 4.8 by 7.283.476.2+7.2
- BrowseCompGemini 3.1 Pro by 1.684.385.9+1.6
ReasoningClaude Opus 4.8
- SimpleBenchGemini 3.1 Pro by 14.864.879.6+14.8
- LiveBench · ReasoningClaude Opus 4.8 by 5.289.284+5.2
- ARC-AGI-2Gemini 3.1 Pro by 572.177.1+5
- CritPtClaude Opus 4.8 by 3.220.917.7+3.2
- GPQA DiamondGemini 3.1 Pro by 2.19294.1+2.1
- Humanity's Last ExamClaude Opus 4.8 by 1.748.747+1.7
- ARC-AGI-3Claude Opus 4.8 by 1.11.50.4+1.1
FactsGemini 3.1 Pro
- SimpleQA VerifiedGemini 3.1 Pro by 20.55373.5+20.5
- AA-Omniscience · Non-hallucinationClaude Opus 4.8 by 11.660.749.1+11.6
- AA-Omniscience · AccuracyGemini 3.1 Pro by 6.148.854.9+6.1
Images and chartsClaude Opus 4.8
- CharXiv (RQ)Claude Opus 4.8 by 6.689.983.3+6.6
- LMArena · Visiontie12951296tie
MathClaude Opus 4.8
- FrontierMath Tier 4Claude Opus 4.8 by 29.356.126.8+29.3
- USAMO 2026Claude Opus 4.8 by 22.396.774.4+22.3
- FrontierMath Tiers 1-3 (v2)Claude Opus 4.8 by 20.38059.6+20.3
- LiveBench · MathematicsClaude Opus 4.8 by 3.394.391+3.3
- IMOAnswerBenchClaude Opus 4.8 by 2.583.581+2.5
- AIME 2026Claude Opus 4.8 by 1.710098.3+1.7
- HMMT Feb. 2026tie95.594.7tie
Following instructionsGemini 3.1 Pro
- IFBenchGemini 3.1 Pro by 14.962.277.1+14.9
- LiveBench · Instruction FollowingGemini 3.1 Pro by 7.17279.1+7.1
Other results46 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDP (Surge AI)Claude Opus 4.8 by 68.184.816.7+68.1
- GDPval-AA v2Claude Opus 4.8 by 631 rating points1593962+631 rating
- GDPval-AA v2 EloClaude Opus 4.8 by 628 rating points1593965+628 rating
- GDPval-AA (Elo)Claude Opus 4.8 by 573 rating points18901317+573 rating
- DeepSWEClaude Opus 4.8 by 485810+48
- DeepSWE 1.1Claude Opus 4.8 by 475912+47
- CyberGymClaude Opus 4.8 by 4078.838.8+40
- NL2RepoClaude Opus 4.8 by 36.369.733.4+36.3
- Program BenchClaude Opus 4.8 by 32.471.939.5+32.4
- AA Agentic IndexClaude Opus 4.8 by 32.342.610.3+32.3
- HLE CalibrationGemini 3.1 Pro by 23.926.550.4+23.9
- SWE-MarathonClaude Opus 4.8 by 22264+22
- PostTrainBenchClaude Opus 4.8 by 15.637.221.6+15.6
- frontiermath_tier_4_v1Claude Opus 4.8 by 14.631.316.7+14.6
- livebench_data_analysisGemini 3.1 Pro by 12.56678.5+12.5
- AA IntelligenceClaude Opus 4.8 by 12.141.829.7+12.1
- Blueprint-Bench 2Gemini 3.1 Pro by 1214.526.5+12
- DeepSearchQA (F1)Claude Opus 4.8 by 11.293.181.9+11.2
- Tool-DecathlonClaude Opus 4.8 by 11.159.948.8+11.1
- ToolathlonClaude Opus 4.8 by 11.159.948.8+11.1
- Finance Agent v2Claude Opus 4.8 by 10.953.943+10.9
- Legal Agent BenchmarkClaude Opus 4.8 by 10100+10
- matharena_visual_math_overallGemini 3.1 Pro by 7.881.689.4+7.8
- AgentWorldBench - TerminalClaude Opus 4.8 by 6.759.252.5+6.7
- CharXiv Reasoning (With tools)Claude Opus 4.8 by 6.789.983.2+6.7
- HLE (with tools)Claude Opus 4.8 by 6.557.951.4+6.5
- OfficeQA ProGemini 3.1 Pro by 6.366.272.5+6.3
- Terminal-Bench 2.0Claude Opus 4.8 by 6.174.668.5+6.1
- Finance AgentGemini 3.1 Pro by 5.853.959.7+5.8
- livebench_languageGemini 3.1 Pro by 5.779.785.4+5.7
- ARC-AGI-1Gemini 3.1 Pro by 5.592.598+5.5
- Artificial Analysis Coding IndexClaude Opus 4.8 by 5.574.368.8+5.5
- AgentWorldBench - SWEClaude Opus 4.8 by 564.159.1+5
- AgentWorldBench - SearchClaude Opus 4.8 by 4.935.130.2+4.9
- AgentWorldBench - MCPGemini 3.1 Pro by 4.254.959.1+4.2
- AA-OmniscienceGemini 3.1 Pro by 3.128.831.9+3.1
- LiveBenchGemini 3.1 Pro by 2.777.279.9+2.7
- AgentWorldBench - OverallClaude Opus 4.8 by 256.654.6+2
- AgentWorldBench - WebClaude Opus 4.8 by 1.954.752.8+1.9
- HMMT Nov. 2025Claude Opus 4.8 by 1.796.594.8+1.7
- τ²-Bench Telecom (AA run)Gemini 3.1 Pro by 1.294.495.6+1.2
- AIRS-BenchClaude Opus 4.8 by 18483+1
- OR-Bench (FRR)tie3.32.5tie
- AgentWorldBench - OStie66.666.9tie
- AgentWorldBench - Androidtie61.561.4tie
- HiL-Bench (Tools-allowed)tie35.335.3tie
Questions people ask
Which is better, Claude Opus 4.8 or Gemini 3.1 Pro?
Claude Opus 4.8 wins five of the eight areas we test: coding, agents, reasoning, images and charts and math. Gemini 3.1 Pro wins facts, long documents and following instructions, and costs 53% less.
Which is better for coding?
Claude Opus 4.8. It wins 8 of the 9 coding tests both models report; Gemini 3.1 Pro wins 1.
Which is cheaper?
Claude Opus 4.8 costs $5.00 per million input tokens and $25.00 per million output tokens; Gemini 3.1 Pro costs $2.00 and $12.00. That makes Gemini 3.1 Pro about 53% cheaper for the same work.
How do you compare the two?
We use the 84 benchmark tests both models have published scores on. The verdict counts the 38 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 46 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.