Gemma 4 31B
Gemma 4 31B is capable in instruction following; and behind the leaders in multimodal tasks, long context, reasoning, and coding. Too few results yet to rate agentic tasks, safety, math, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Math, Multilingual or Factuality.
Price
$0.17input$0.40outputper million tokens
From Artificial Analysis · 4 providers tracked · All prices
Evidence
111results on81benchmarks
- 10 independently verified
- 38 aggregator
- 17 vendor-reported
- 46 cross-referenced
From 11 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingJSON modeReasoning
As listed by OpenRouter
Research
31 papers reference Gemma 4 31BGemma 4 31B benchmark results
111 results on 81 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
24.0% behind the leader1 of 3 ranked benchmarks measured
- IFBench75.58Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench53.47Oct 8, 2026aa_ifbench
25.6% behind the leader5 of 6 ranked benchmarks measured
- 80.40Aug 24, 2026
- 86.10Aug 24, 2026
- MathVista79.30Aug 24, 2026Mathvista(mini)
- MMMU-Pro73.41Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)67.90Aug 24, 2026CharXiv(RQ)
Show 6 more multimodal resultsHide 6 multimodal results
- 1276.92Aug 25, 2026
- 1252Jun 17, 2026
- MMMU-Pro70.35Oct 8, 2026aa_mmmu_pro
- 76.90Oct 8, 2026
- 76.90Aug 24, 2026
- 77.30Aug 24, 2026
26.9% behind the leader2 of 3 ranked benchmarks measured
- 69.67Oct 8, 2026
- MRCR v2 (8-needle, 128K)66.40Oct 8, 2026MRCR v2 8 needle 128k (average)
Show 1 more long context resultHide 1 long context result
- 46.67Oct 8, 2026
31.8% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond85.66Oct 8, 2026gpqa
- Humanity's Last Exam23.63Oct 8, 2026aa_hle
- 1.43Oct 8, 2026
Show 8 more reasoning resultsHide 8 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond76.26Oct 8, 2026gpqa
- 84.30Oct 8, 2026
- GPQA Diamond84.30Aug 24, 2026GPQA
- Humanity's Last Exam11.82Oct 8, 2026aa_hle
- Humanity's Last Exam19.50Oct 8, 2026HLE no tools
- 26.50Oct 7, 2026
- 19.50Jun 15, 2026
39.3% behind the leader8 of 10 ranked benchmarks measured
- 80.00Oct 8, 2026
- SciCode45.49Oct 8, 2026aa_scicode
- Terminal-Bench Hard36.36Oct 8, 2026aa_terminalbench_hard
- 51.70Jun 15, 2026
- 52.00Jun 15, 2026
- 1364.68May 22, 2026
- Terminal-Bench 2.143.45Oct 8, 2026terminalbenchV21
- 35.70Jun 15, 2026
Show 4 more coding resultsHide 4 coding results
- 80.00Jun 15, 2026
- SciCode41.09Sep 4, 2026aa_scicode
- Terminal-Bench 2.129.21Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard30.30Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 37.29Oct 8, 2026
- 2.52Oct 8, 2026
- 6.15Oct 8, 2026
Show 4 more agentic resultsHide 4 agentic results
- 57.20Jun 6, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking14.85Oct 8, 2026tauBanking
- τ-Bench V3 · Banking8.87Oct 8, 2026tauBanking
0 of 5 ranked benchmarks measured
Show 1 more math resultHide 1 math result
- HMMT Feb 202677.20Jun 15, 2026HMMT Feb 26
0 of 4 ranked benchmarks measured
- 10.40May 20, 2026
- 7.40May 2, 2026
- AA-Omniscience · Accuracy16.63Oct 8, 2026omniscienceAccuracy
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy20.03Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination15.01Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination18.05Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1451Jul 23, 2026
- AA Intelligence29.00Jul 3, 2026Artificial Analysis Intelligence Index
- vectara_avg_summary_length75.80May 2, 2026Average Summary Length (Words)
- vectara_answer_rate100.00May 2, 2026Answer Rate
- vectara_factual_consistency92.60May 2, 2026Factual Consistency Rate
- AA Intelligence14.67Oct 8, 2026aa_intelligence_index
Show 49 more resultsHide 49 results
- 6.73Sep 9, 2026
- 11.10Sep 4, 2026
- AA Intelligence13.92Oct 8, 2026aa_intelligence_index
- -51.68Oct 8, 2026
- -47.93Oct 8, 2026
- AI2D89.00Aug 24, 2026AI2D_TEST
- Artificial Analysis Coding Index33.17Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index43.43Sep 9, 2026aa_coding_index
- 74.40Oct 8, 2026
- 82.60Jun 15, 2026
- 75.70Aug 24, 2026
- Claw Eval (pass@3)25.00Aug 24, 2026Claw-Eval Pass^3
- 48.50Jun 15, 2026
- 2150.00Oct 8, 2026
- 24.00Jun 6, 2026
- 79.50Aug 24, 2026
- 57.50Aug 24, 2026
- 84.30Jun 6, 2026
- 67.40Aug 24, 2026
- HLE (with tools)26.50Oct 8, 2026HLE with search
- HMMT Feb. 202588.70Jun 15, 2026HMMT Feb 25
- HMMT Nov. 202587.50Jun 15, 2026HMMT Nov 25
- 85.60Oct 8, 2026
- 18.10Jun 6, 2026
- 61.30Oct 8, 2026
- 90.90Aug 24, 2026
- 85.20Oct 8, 2026
- 85.20Jun 15, 2026
- 93.70Jun 15, 2026
- 88.40Oct 8, 2026
- 15.50Jun 15, 2026
- 80.10Aug 24, 2026
- OmniDocBench 1.5 (average edit distance, lower is better)lower is better0.13Oct 8, 2026
- 41.70Jun 15, 2026
- 1197.00Jun 15, 2026
- 72.30Aug 24, 2026
- 4.70Aug 24, 2026
- 52.90Aug 24, 2026
- 23.60Jun 15, 2026
- 65.70Jun 15, 2026
- 86.40Oct 7, 2026
- 42.90Jun 15, 2026
- 21.20Jun 6, 2026
- 43.00Jun 6, 2026
- 35.20Jun 6, 2026
- τ²-Bench76.90Oct 8, 2026Tau2 (average over 3)
- τ²-Bench Telecom (AA run)59.94Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)65.50Oct 8, 2026aa_tau2
- τ³-Bench67.50Jun 6, 2026TAU3-Bench
Gemma 4 31B: common questions
Who makes Gemma 4 31B?
Gemma 4 31B is made by Google.
When was Gemma 4 31B released?
Gemma 4 31B was released on Apr 2, 2026, according to Artificial Analysis.
What is Gemma 4 31B good at?
Gemma 4 31B is capable in instruction following; and behind the leaders in multimodal tasks, long context, reasoning, and coding. Too few results yet to rate agentic tasks, safety, math, multilingual tasks, or factuality.
How much does Gemma 4 31B cost?
Gemma 4 31B costs $0.17 per million input tokens and $0.40 per million output tokens, according to Artificial Analysis. We track its price at 4 providers. At a mix of three input tokens to one output token, it is cheaper than 79% of the 331 priced models we track.
How many benchmarks has Gemma 4 31B been tested on?
We track 111 results for Gemma 4 31B on 81 benchmarks from 11 sources, 10 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Gemma 4 31B support?
OpenRouter lists tool calling, json mode, and reasoning for Gemma 4 31B.
About this record
Where Gemma 4 31B's numbers come from, and every name it appears under.
- Tracked since
- Apr 30, 2026
- Newest source mention
- Sep 4, 2026
Where the results come from
Verification: 111 scores · 10 independently verified · 38 aggregator-attributed · 46 vendor cross-reference · 17 vendor-reported. How these tiers are assigned
From 11 sources on 7 sites. Hugging Face supplies 61 of them; the 10 independently verified results come from 5 sites. Bars are coloured by trust tier.
- huggingface.co61
- artificialanalysis.ai39
- raw.githubusercontent.com4
- api.llm-stats.com2
- datasets-server.huggingface.co2
- lmarena.ai2
- epoch.ai1