Gemma 4 E4B
Gemma 4 E4B is behind the leaders in factuality, instruction following, and multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math or Multilingual.
Price
$0.02input$0.10outputper million tokens
From Artificial Analysis · All prices
Evidence
118results on73benchmarks
- 1 independently verified
- 34 aggregator
- 52 vendor-reported
- 31 cross-referenced
From 9 sources · latest Oct 8, 2026 · How verification works
Research
20 papers reference Gemma 4 E4BGemma 4 E4B benchmark results
118 results on 73 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
37.6% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Non-hallucination69.06Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Accuracy8.58Oct 8, 2026omniscienceAccuracy
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy8.32Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination46.23Oct 8, 2026omniscienceNonHallucination
41.6% behind the leader1 of 3 ranked benchmarks measured
- IFBench44.22Oct 8, 2026aa_ifbench
47.8% behind the leader2 of 6 ranked benchmarks measured
Show 10 more multimodal resultsHide 10 multimodal results
- 52.20Sep 10, 2026
- 56.40Sep 25, 2026
- 52.90Sep 25, 2026
- MMMU-Pro51.21Oct 8, 2026aa_mmmu_pro
- 52.60Oct 8, 2026
- 32.60Sep 10, 2026
- 39.10Sep 25, 2026
- 52.90Sep 10, 2026
- 61.90Sep 25, 2026
- 48.70Sep 25, 2026
0 of 6 ranked benchmarks measured
- 0.57Oct 8, 2026
- GPQA Diamond57.58Oct 8, 2026gpqa
- Humanity's Last Exam3.75Oct 8, 2026aa_hle
Show 4 more reasoning resultsHide 4 reasoning results
- 0.29Oct 8, 2026
- GPQA Diamond54.95Oct 8, 2026gpqa
- 58.60Oct 8, 2026
- Humanity's Last Exam4.82Oct 8, 2026aa_hle
0 of 10 ranked benchmarks measured
- SciCode24.42Oct 8, 2026aa_scicode
- Terminal-Bench 2.11.87Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard8.33Oct 8, 2026aa_terminalbench_hard
Show 3 more coding resultsHide 3 coding results
- 52.00Oct 8, 2026
- SciCode3.94Sep 4, 2026aa_scicode
- Terminal-Bench Hard7.58Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking5.36Aug 10, 2026tauBanking
0 of 3 ranked benchmarks measured
- 32.00Oct 8, 2026
- 24.00Oct 8, 2026
- MRCR v2 (8-needle, 128K)25.40Oct 8, 2026MRCR v2 8 needle 128k (average)
0 of 5 ranked benchmarks measured
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence12.00Jul 11, 2026Artificial Analysis Intelligence Index
- AA Intelligence8.91Oct 8, 2026aa_intelligence_index
- -19.70Oct 8, 2026
- τ²-Bench Telecom (AA run)20.76Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)26.02Oct 8, 2026aa_tau2
- -40.98Oct 8, 2026
Show 69 more resultsHide 69 results
- 1.79Aug 10, 2026
- 8.67Jun 18, 2026
- AA Intelligence7.46Oct 8, 2026aa_intelligence_index
- AIME25 no tools34.27Aug 10, 2026AIME25
- AIME25 no tools34.33Aug 9, 2026AIME25
- AIME25 no tools34.27Aug 9, 2026AIME25
- Artificial Analysis Coding Index9.39Sep 9, 2026aa_coding_index
- 6.36Jun 18, 2026
- 57.31Aug 9, 2026
- 40.00Sep 10, 2026
- 46.39Aug 10, 2026
- 40.00Sep 25, 2026
- 33.92Aug 9, 2026
- 46.39Aug 9, 2026
- 33.10Oct 8, 2026
- 15.90Aug 10, 2026
- 15.90Aug 9, 2026
- 41.90Sep 25, 2026
- ChartQA Test42.10Sep 10, 2026ChartQA (test)
- Claw-Eval Avg58.02Aug 10, 2026Claw-Eval average (EN)
- Claw-Eval Avg58.02Aug 9, 2026Claw-Eval average (EN)
- 940.00Oct 8, 2026
- 35.54Oct 8, 2026
- DocVQA-val87.40Sep 10, 2026DocVQA (val)
- FLEURS (lower is better)lower is better0.08Oct 8, 2026
- 49.80Sep 10, 2026
- 87.90Sep 10, 2026
- 87.74Aug 9, 2026
- 60.90Sep 10, 2026
- LiveCodeBenchV6 no tools63.77Aug 10, 2026LiveCodeBenchv6
- LiveCodeBenchV6 no tools63.77Aug 9, 2026LiveCodeBenchv6
- 59.50Oct 8, 2026
- MATH500 (Pass@1)65.00Aug 9, 2026MATH500
- 28.70Oct 8, 2026
- 68.20Sep 10, 2026
- 66.70Sep 25, 2026
- MMBench71.60Sep 10, 2026MMBench (dev EN v1.1)
- 67.60Sep 10, 2026
- 68.10Sep 25, 2026
- 69.40Oct 8, 2026
- 80.40Sep 19, 2026
- 80.50Sep 25, 2026
- 76.60Oct 8, 2026
- MMMU (val) (Pass@1)49.30Sep 10, 2026MMMU (val)
- 71.20Sep 10, 2026
- 80.40Sep 10, 2026
- 73.50Sep 10, 2026
- OCRBench v2_en48.80Sep 10, 2026OCRBench v2 (En)
- OmniDocBench 1.5 (average edit distance, lower is better)lower is better0.18Oct 8, 2026
- 55.09Aug 10, 2026
- 55.09Aug 9, 2026
- 86.90Sep 10, 2026
- 86.90Sep 25, 2026
- 64.30Sep 10, 2026
- 61.80Sep 25, 2026
- 72.10Sep 25, 2026
- 50.90Sep 25, 2026
- 45.80Sep 10, 2026
- 60.30Sep 10, 2026
- 47.60Sep 10, 2026
- 75.30Sep 10, 2026
- 30.40Sep 10, 2026
- 57.50Oct 7, 2026
- 26.75Aug 9, 2026
- TextVQA-val69.00Sep 10, 2026TextVQA (val)
- τ²-Bench42.20Oct 8, 2026Tau2 (average over 3)
- τ²-Bench (Retail)42.11Aug 9, 2026Tau² Retail
- 4.12Aug 10, 2026
- 4.12Aug 9, 2026
Gemma 4 E4B: common questions
Who makes Gemma 4 E4B?
Gemma 4 E4B is made by Google.
When was Gemma 4 E4B released?
Gemma 4 E4B was released on Apr 3, 2026, according to Artificial Analysis.
What is Gemma 4 E4B good at?
Gemma 4 E4B is behind the leaders in factuality, instruction following, and multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, or multilingual tasks.
How much does Gemma 4 E4B cost?
Gemma 4 E4B costs $0.02 per million input tokens and $0.10 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it is cheaper than 98% of the 330 priced models we track.
How many benchmarks has Gemma 4 E4B been tested on?
We track 118 results for Gemma 4 E4B on 73 benchmarks from 9 sources, 1 of them independently verified. The latest was recorded on Oct 8, 2026.
About this record
Where Gemma 4 E4B's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Sep 25, 2026
Where the results come from
Verification: 118 scores · 1 independently verified · 34 aggregator-attributed · 31 vendor cross-reference · 52 vendor-reported. How these tiers are assigned
From 9 sources on 3 sites. Hugging Face supplies 82 of them; the 1 independently verified result comes from 1 site. Bars are coloured by trust tier.
- huggingface.co82
- artificialanalysis.ai35
- api.llm-stats.com1