Gemma 3 4B vs Llama 3.2 3B Instruct
Wins 4 of 5 areas
Coding · Reasoning · Long documents · Following instructions
Wins 0 of 5 areas
—
Gemma 3 4B is the stronger all-rounder.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20Gemma 3 4B2 of 2 tests
- ReasoningHard problems that need careful thinking10Gemma 3 4B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts10Gemma 3 4B1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemma 3 4B1 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemma 3 4B costs 61% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemma 3 4B pulls ahead
- Recent programming contest problemsLiveCodeBench v6+7.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+3.6points ahead
- Reasons across sets of long documentsAA-LCR+2.4points ahead
Where Llama 3.2 3B Instruct pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+40.7points ahead
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemma 3 4B
- LiveCodeBench v6Gemma 3 4B by 7.418.911.5+7.4
- SciCodeGemma 3 4B by 2.17.35.2+2.1
ReasoningGemma 3 4B
- GPQA DiamondGemma 3 4B by 3.629.125.5+3.6
- Humanity's Last Examtie5.35.3tie
FactsEven
- AA-Omniscience · Non-hallucinationLlama 3.2 3B Instruct by 40.71.842.5+40.7
- AA-Omniscience · AccuracyGemma 3 4B by 1.47.76.3+1.4
Following instructionsGemma 3 4B
- IFBenchGemma 3 4B by 2.128.326.2+2.1
Other results16 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HumanEvalGemma 3 4B by 43.371.328+43.3
- HumanEval+Gemma 3 4B by 38.762.824.1+38.7
- AA-OmniscienceLlama 3.2 3B Instruct by 35.4-82.9-47.5+35.4
- MATH-500 (EM)Gemma 3 4B by 3273.241.2+32
- GSMPlusGemma 3 4B by 29.768.438.7+29.7
- MGSMGemma 3 4B by 29.187.358.2+29.1
- BBHGemma 3 4B by 25.472.246.8+25.4
- τ²-Bench Telecom (AA run)Llama 3.2 3B Instruct by 16.1521.1+16.1
- MBPPGemma 3 4B by 14.563.248.7+14.5
- IFEvalGemma 3 4B by 12.890.277.4+12.8
- GSM8KGemma 3 4B by 11.589.277.7+11.5
- MMLU-ProGemma 3 4B by 8.643.635+8.6
- LCB v5Gemma 3 4B by 7.619.111.5+7.6
- MMMLUGemma 3 4B by 2.250.147.9+2.2
- MMLULlama 3.2 3B Instruct by 258.460.4+2
- AA Intelligencetie4.85.7tie
Questions people ask
Which is better, Gemma 3 4B or Llama 3.2 3B Instruct?
Gemma 3 4B wins four of the five areas where both have results: coding, reasoning, long documents and following instructions. Llama 3.2 3B Instruct wins none. They are level on facts.
Which is better for coding?
Gemma 3 4B. It wins 2 of the 2 coding tests both models report; Llama 3.2 3B Instruct wins none.
Which is cheaper?
Gemma 3 4B costs $0.05 per million input tokens and $0.10 per million output tokens; Llama 3.2 3B Instruct costs $0.05 and $0.33. That makes Gemma 3 4B about 61% cheaper for the same work.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 16 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.