Gemma 4 12B
Gemma 4 12B is behind the leaders in instruction following, multimodal tasks, long context, and coding. Too few results yet to rate reasoning, agentic tasks, safety, math, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Agentic, Safety, Math, Multilingual or Factuality.
Price
$0.10input$0.30outputper million tokens
From Artificial Analysis · All prices
Evidence
55results on35benchmarks
- 4 independently verified
- 33 aggregator
- 18 vendor-reported
From 8 sources · latest Oct 8, 2026 · How verification works
Research
3 papers reference Gemma 4 12BGemma 4 12B benchmark results
55 results on 35 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
25.5% behind the leader1 of 3 ranked benchmarks measured
- IFBench73.54Oct 8, 2026aa_ifbench
Show 1 more instruction following resultHide 1 instruction following result
- IFBench45.17Oct 8, 2026aa_ifbench
32.2% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro69.65Oct 8, 2026aa_mmmu_pro
Show 3 more multimodal resultsHide 3 multimodal results
- MMMU-Pro61.97Oct 8, 2026aa_mmmu_pro
- 69.10Oct 8, 2026
- 51.80Jul 2, 2026
33.7% behind the leader2 of 3 ranked benchmarks measured
- 63.67Oct 8, 2026
- MRCR v2 (8-needle, 128K)43.40Oct 8, 2026MRCR v2 8 needle 128k (average)
Show 1 more long context resultHide 1 long context result
- 35.00Oct 8, 2026
37.7% behind the leader4 of 10 ranked benchmarks measured
- 72.00Oct 8, 2026
- SciCode38.19Sep 4, 2026aa_scicode
- Terminal-Bench 2.127.34Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard18.18Oct 8, 2026aa_terminalbench_hard
Show 2 more coding resultsHide 2 coding results
- SciCode29.75Sep 4, 2026aa_scicode
- Terminal-Bench Hard11.36Oct 8, 2026aa_terminalbench_hard
0 of 6 ranked benchmarks measured
- 0.00Oct 8, 2026
- GPQA Diamond75.25Oct 8, 2026gpqa
- Humanity's Last Exam15.66Oct 8, 2026aa_hle
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond66.06Oct 8, 2026gpqa
- 78.80Oct 8, 2026
- Humanity's Last Exam6.35Oct 8, 2026aa_hle
- 5.20Oct 7, 2026
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- τ-Bench V3 · Banking8.66Aug 10, 2026tauBanking
- 2.69Jun 15, 2026
0 of 5 ranked benchmarks measured
- AIME 202677.50Oct 8, 2026AIME 2026 no tools
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy15.62Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination19.02Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy11.88Oct 8, 2026omniscienceAccuracy
Show 1 more factuality resultHide 1 factuality result
- AA-Omniscience · Non-hallucination26.59Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 37.50Sep 11, 2026
- AA Intelligence22.00Jul 5, 2026Artificial Analysis Intelligence Index
- AA Intelligence9.00Jun 11, 2026Artificial Analysis Intelligence Index
- -52.72Oct 8, 2026
- AA Intelligence14.18Oct 8, 2026aa_intelligence_index
- τ²-Bench Telecom (AA run)36.26Oct 8, 2026aa_tau2
Show 19 more resultsHide 19 results
- 7.93Aug 10, 2026
- 12.38Jun 18, 2026
- AA Intelligence9.38Oct 8, 2026aa_intelligence_index
- -52.80Oct 8, 2026
- Artificial Analysis Coding Index30.96Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index17.49Jun 18, 2026aa_coding_index
- 53.00Oct 8, 2026
- 1659.00Oct 8, 2026
- 38.50Oct 8, 2026
- 93.10Oct 7, 2026
- FLEURS (lower is better)lower is better0.07Oct 8, 2026
- HLE (with tools)5.20Jul 16, 2026HLE with search
- 79.70Oct 8, 2026
- 48.70Oct 8, 2026
- 77.20Oct 8, 2026
- 83.40Oct 8, 2026
- OmniDocBench 1.5 (average edit distance, lower is better)lower is better0.16Oct 8, 2026
- τ²-Bench69.00Oct 8, 2026Tau2 (average over 3)
- τ²-Bench Telecom (AA run)31.87Oct 8, 2026aa_tau2
Gemma 4 12B: common questions
Who makes Gemma 4 12B?
Gemma 4 12B is made by Google.
When was Gemma 4 12B released?
Gemma 4 12B was released on Jun 3, 2026, according to Artificial Analysis.
What is Gemma 4 12B good at?
Gemma 4 12B is behind the leaders in instruction following, multimodal tasks, long context, and coding. Too few results yet to rate reasoning, agentic tasks, safety, math, multilingual tasks, or factuality.
How much does Gemma 4 12B cost?
Gemma 4 12B costs $0.10 per million input tokens and $0.30 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it is cheaper than 85% of the 330 priced models we track.
How many benchmarks has Gemma 4 12B been tested on?
We track 55 results for Gemma 4 12B on 35 benchmarks from 8 sources, 4 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Gemma 4 12B's numbers come from, and every name it appears under.
- Tracked since
- Jun 18, 2026
- Newest source mention
- Sep 7, 2026
Where the results come from
Verification: 55 scores · 4 independently verified · 33 aggregator-attributed · 18 vendor-reported. How these tiers are assigned
From 8 sources on 4 sites. Artificial Analysis supplies 35 of them; the 4 independently verified results come from 2 sites. Bars are coloured by trust tier.
- artificialanalysis.ai35
- huggingface.co16
- 99franklin.github.io2
- api.llm-stats.com2