Gemini 3.1 Pro
Gemini 3.1 Pro is at the frontier in instruction following; strong in factuality, multimodal tasks, and reasoning; capable in long context and coding; and behind the leaders in agentic tasks and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$2.00input$12.00outputper million tokens
From Artificial Analysis · 4 providers tracked · All prices
Evidence
217results on174benchmarks
- 30 independently verified
- 21 aggregator
- 37 vendor-reported
- 129 cross-referenced
From 37 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Gemini 3.1 Pro benchmark results
217 results on 174 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
Leads the field3 of 3 ranked benchmarks measured
- LiveBench · Instruction Following79.10Oct 8, 2026livebench_instruction_following@2026-06-25
- IFBench77.14Oct 8, 2026aa_ifbench
- 71.37Oct 8, 2026
Show 1 more instruction following resultHide 1 instruction following result
- 77.10Jul 16, 2026
2.8% behind the leader4 of 4 ranked benchmarks measured
- SimpleQA Verified75.60Jun 27, 2026SimpleQA-Verified (Pass@1)
- 10.40May 2, 2026
- AA-Omniscience · Accuracy54.85Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination49.13Oct 8, 2026omniscienceNonHallucination
Show 1 more factuality resultHide 1 factuality result
- 77.30Jul 16, 2026
8.6% behind the leader4 of 6 ranked benchmarks measured
- 90.20Sep 28, 2026
- 86.70Sep 28, 2026
- MMMU-Pro82.43Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)83.30Jul 6, 2026CharXiv Reasoning Information synthesis from complex charts
Show 7 more multimodal resultsHide 7 multimodal results
- 79.10Sep 28, 2026
- CharXiv (reasoning)80.20May 3, 2026CharXiv (RQ)
- 1295.82Aug 25, 2026
- 1280Jun 17, 2026
- 80.50Oct 7, 2026
- 80.50Sep 28, 2026
- 62.80Sep 28, 2026
9.2% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond94.14Oct 8, 2026gpqa
- LiveBench · Reasoning84.00Oct 8, 2026livebench_reasoning@2026-06-25
- 79.60May 10, 2026
- ARC-AGI-277.10Oct 7, 2026ARC-AGI v2
- Humanity's Last Exam47.03Oct 8, 2026aa_hle
- 17.71Oct 8, 2026
Show 7 more reasoning resultsHide 7 reasoning results
- 0.42May 10, 2026
- 17.70Jul 13, 2026
- GPQA Diamond94.30Oct 7, 2026GPQA
- 94.30Oct 8, 2026
- 51.40Oct 7, 2026
- 45.00Oct 8, 2026
- Humanity's Last Exam44.40Jun 27, 2026HLE (Pass@1)
13.5% behind the leader2 of 3 ranked benchmarks measured
- 82.00Oct 8, 2026
- MRCR v2 (8-needle, 128K)26.30Aug 24, 2026MRCR v2 (8-needle)
Show 1 more long context resultHide 1 long context result
- MRCR v2 (8-needle, 128K)84.90Jul 6, 2026MRCR v2 (8-needle) Long context performance
20.4% behind the leader10 of 10 ranked benchmarks measured
- LiveCodeBench v691.70May 3, 2026LiveCodeBench (v6)
- SciCode58.68Oct 8, 2026aa_scicode
- 80.60Oct 7, 2026
- LiveBench · Coding76.45Oct 8, 2026livebench_coding@2026-06-25
- 76.90May 3, 2026
- Terminal-Bench Hard53.79Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.173.78Oct 8, 2026terminalbenchV21
- 1446.57May 22, 2026
- 54.20Oct 7, 2026
- LiveBench · Agentic Coding44.14Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 7 more coding resultsHide 7 coding results
- 59.00Oct 7, 2026
- 46.10Oct 8, 2026
- 54.20Oct 8, 2026
- 80.60Jul 16, 2026
- Terminal-Bench 2.170.70Jul 13, 2026Terminal Bench 2.1 (Best Reported Harness)
- Terminal-Bench 2.174.00Jul 13, 2026Terminal Bench 2.1 (Terminus-2)
- 68.50May 1, 2026
28.4% behind the leader7 of 7 ranked benchmarks measured
- 85.90Oct 7, 2026
- OSWorld-Verified76.20Jul 6, 2026OSWorld-Verified Agentic computer use
- 69.20Oct 7, 2026
- τ-Bench V3 · Banking21.44Oct 8, 2026tauBanking
- 14.72Oct 8, 2026
- Terminal-Bench 4.04.04Oct 8, 2026
Show 9 more agentic resultsHide 9 agentic results
- 32.01Oct 8, 2026
- AA ApexAgents33.50Oct 7, 2026APEX-Agents
- AA ApexAgents33.50Sep 28, 2026APEX Agents
- 30.33Oct 8, 2026
- BrowseComp85.90Jul 4, 2026BrowseComp (Pass@1)
- GDPval (win rate)67.30Sep 28, 2026GDPval
- 78.20Oct 8, 2026
- MCP Atlas69.20Oct 8, 2026MCP-Atlas (Public Set)
- 76.20Jun 12, 2026
40.0% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics91.04Oct 8, 2026livebench_math@2026-06-25
- 59.65Sep 9, 2026
- 26.83Sep 8, 2026
Show 7 more math resultsHide 7 math results
- 98.33Sep 2, 2026
- 98.20Oct 8, 2026
- 98.30Jul 16, 2026
- 94.70Sep 2, 2026
- HMMT Feb 202687.30Oct 8, 2026HMMT Feb. 2026
- 81.00Oct 8, 2026
- 74.40Sep 2, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1487Oct 8, 2026
- livebench_data_analysis78.54Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language85.38Oct 8, 2026livebench_language@2026-06-25
- 35.33Oct 8, 2026
- 89.44Sep 2, 2026
- AA Intelligence30.00Jul 29, 2026Artificial Analysis Intelligence Index
Show 133 more resultsHide 133 results
- 10.31Sep 9, 2026
- AA Intelligence29.72Oct 8, 2026aa_intelligence_index
- 31.88Oct 8, 2026
- 61.40Jun 24, 2026
- 59.07Jun 24, 2026
- 66.92Jun 24, 2026
- 54.57Jun 24, 2026
- 30.21Jun 24, 2026
- 59.07Jun 24, 2026
- 52.47Jun 24, 2026
- 52.83Jun 24, 2026
- 83.00Jul 13, 2026
- 60.90Jul 4, 2026
- 89.10Jul 4, 2026
- 98.00May 10, 2026
- 98.00Jun 21, 2026
- Artificial Analysis Coding Index68.83Sep 9, 2026aa_coding_index
- baby_vision_with_python68.30May 3, 2026BabyVision (w/ python)
- 54.40Sep 28, 2026
- 51.60May 3, 2026
- Blueprint-Bench 226.50Jul 6, 2026Blueprint-Bench 2 Agentic spatial reasoning
- 26.50Jun 12, 2026
- 85.90Jun 6, 2026
- 85.90Jul 16, 2026
- browsecomp_with_context_manager85.90Oct 8, 2026BrowseComp (w/ Context Manage)
- 72.50Sep 2, 2026
- 70.20Sep 28, 2026
- 83.20Jul 22, 2026
- charxiv_rq_with_python89.90Jul 16, 2026Charxiv RQ (with python)
- Chinese SimpleQA (C-SimpleQA)85.90Jul 4, 2026Chinese-SimpleQA (Pass@1)
- 82.90May 3, 2026
- 3052.00Jul 4, 2026
- 53.80Jul 4, 2026
- 38.80Oct 8, 2026
- 60.20May 3, 2026
- DeepSearchQA (F1)81.90May 3, 2026DeepSearchQA (f1-score)
- 10.00Sep 28, 2026
- 12.00Oct 7, 2026
- 82.10Sep 28, 2026
- 72.10Sep 28, 2026
- 84.20Sep 28, 2026
- 70.80Sep 28, 2026
- 59.70Jun 21, 2026
- Finance Agent43.00Jun 1, 2026Finance Agent v2
- Finance Agent v1.159.70Sep 28, 2026
- 42.98Oct 7, 2026
- Finance Agent v243.00Jul 6, 2026Finance Agent v2 Financial analysis and decision-making
- 61.90Jul 30, 2026
- frontiermath_tier_4_v116.70May 20, 2026frontiermath_tier_4
- 26.30Jul 22, 2026
- GDP (Surge AI)16.70Jun 12, 2026GDP.pdf
- GDPval-AA (Elo)1317.00Aug 24, 2026GDPval-AA Elo
- GDPval-AA (Elo)1314.00Jul 6, 2026GDPval-AA Economically valuable knowledge work
- 1314.00Jun 27, 2026
- 962.00Jul 16, 2026
- GDPval-AA v2 Elo965.00Jul 22, 2026GDPVal-AA v2 (Elo)
- 92.70Jul 16, 2026
- HLE (with tools)51.40Oct 8, 2026HLE (w/ Tools)
- 50.40Jul 13, 2026
- 94.80Oct 8, 2026
- 44.40Jul 7, 2026
- Key Information Extraction Overall79.20Sep 2, 2026
- 0.00Oct 7, 2026
- 79.93Oct 7, 2026
- LiveCodeBench91.70Jul 4, 2026LiveCodeBench (Pass@1)
- 2887.00Aug 24, 2026
- 2887.00May 1, 2026
- 76.50Sep 28, 2026
- 66.20Sep 28, 2026
- 89.80May 3, 2026
- MathVerse (Vision-Only)87.70Sep 28, 2026
- mathvision_with_python95.70May 3, 2026MathVision (w/ python)
- 69.20Jul 4, 2026
- 55.90May 3, 2026
- 63.50Sep 28, 2026
- 42.60Jul 22, 2026
- 82.50Jul 16, 2026
- 70.70Sep 28, 2026
- 92.60Aug 24, 2026
- MMLU-Pro91.00Jul 4, 2026MMLU-Pro (EM)
- 92.60Oct 7, 2026
- 92.60Jun 21, 2026
- mmmu_pro_with_python85.30May 3, 2026MMMU-Pro (w/ python)
- MMSIBench (circular)27.40Sep 28, 2026
- 69.90Sep 28, 2026
- 26.30May 1, 2026
- 33.40Oct 8, 2026
- OCRBench KIE96.00Sep 2, 2026
- OCRBenchv2 KIE (en)87.80Sep 2, 2026
- OCRBenchv2 KIE (zh)63.40Sep 2, 2026
- Office QA Pro [Multimodal]72.50Sep 28, 2026
- 72.50Sep 28, 2026
- 42.90Jun 21, 2026
- 18.10Jun 12, 2026
- 70.70May 3, 2026
- OpenAI-MRCR (1M)76.30Jul 4, 2026MRCR 1M (MMR)
- 2.50Jul 13, 2026
- 58.80Sep 28, 2026
- 21.60Jul 13, 2026
- 39.50Jul 13, 2026
- 0.12Jul 30, 2026
- 37.00Sep 1, 2026
- 85.40Sep 28, 2026
- 69.90Sep 28, 2026
- 98.00Jul 16, 2026
- 4.00Jul 13, 2026
- SWE-Pro Bench54.20Sep 28, 2026
- SWEBench Pro Public54.20Jul 16, 2026SWEBench Pro (Public)
- 99.30Oct 7, 2026
- 68.50Oct 7, 2026
- Terminal-Bench 2.068.50Oct 8, 2026Terminal-Bench 2.0 (Terminus-2)
- 60.40Sep 28, 2026
- 48.80Oct 8, 2026
- 48.80Sep 28, 2026
- 71.00Sep 28, 2026
- 96.90May 3, 2026
- vectara_answer_rate99.40May 2, 2026Answer Rate
- vectara_avg_summary_length107.70May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency89.60May 2, 2026Factual Consistency Rate
- 911.21Oct 8, 2026
- 65.90Sep 28, 2026
- 70.00Sep 28, 2026
- 44.80Sep 28, 2026
- 94.30Jul 16, 2026
- WMDP86.30Jul 13, 2026WMDP-Chem
- 44.30Sep 28, 2026
- ZeroBench (main)12.00Sep 28, 2026
- ZeroBench (sub)41.90Sep 28, 2026
- 90.80May 6, 2026
- τ²-Bench Telecom (AA run)95.61Oct 8, 2026aa_tau2
- 99.30May 1, 2026
- 67.10Oct 8, 2026
- τ³-Bench Banking16.50Jul 16, 2026Tau 3 Banking
Gemini 3.1 Pro: common questions
Who makes Gemini 3.1 Pro?
Gemini 3.1 Pro is made by Google.
When was Gemini 3.1 Pro released?
Gemini 3.1 Pro was released on Feb 19, 2026, according to Artificial Analysis.
What is Gemini 3.1 Pro good at?
Gemini 3.1 Pro is at the frontier in instruction following; strong in factuality, multimodal tasks, and reasoning; capable in long context and coding; and behind the leaders in agentic tasks and math. Too few results yet to rate safety or multilingual tasks.
How much does Gemini 3.1 Pro cost?
Gemini 3.1 Pro costs $2.00 per million input tokens and $12.00 per million output tokens, according to Artificial Analysis. We track its price at 4 providers. At a mix of three input tokens to one output token, it costs more than 85% of the 330 priced models we track.
How many benchmarks has Gemini 3.1 Pro been tested on?
We track 217 results for Gemini 3.1 Pro on 174 benchmarks from 37 sources, 30 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Gemini 3.1 Pro support?
OpenRouter lists tool calling, structured outputs, and reasoning for Gemini 3.1 Pro.
About this record
Where Gemini 3.1 Pro's numbers come from, and every name it appears under.
- Tracked since
- Apr 26, 2026
- Newest source mention
- Sep 28, 2026
Where the results come from
Verification: 217 scores · 30 independently verified · 21 aggregator-attributed · 129 vendor cross-reference · 37 vendor-reported. How these tiers are assigned
From 37 sources on 19 sites. Hugging Face supplies 76 of them; the 30 independently verified results come from 10 sites. Bars are coloured by trust tier.
- huggingface.co76
- lf3-static.bytednsdoc.com37
- artificialanalysis.ai22
- api.llm-stats.com17
- deepmind.google16
- www-cdn.anthropic.com9
- livebench.ai7
- ai.meta.com4
- labs.scale.com4
- matharena.ai4
- raw.githubusercontent.com4
- storage.googleapis.com4
- epoch.ai3
- arcprize.org2
- datasets-server.huggingface.co2
- lmarena.ai2
- thinkingmachines.ai2
- cdn.sanity.io1
- simple-bench.com1