Granite 4.0 1B wins more areas, narrowly.
Scores updated · 26 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking01Granite 4.0 1B1 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts01Granite 4.0 1B1 of 1 test
- CodingWriting and fixing software00Even0 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
- Following instructionsDoing exactly what it is asked00Even0 each · 1 tie
Coding, long documents and following instructions rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Granite 4.0 1B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4.4points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+2.4points ahead
Where Gemma 3 1B Instruct pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+11.2points ahead
Every test, side by side
All 26 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningGranite 4.0 1B
- GPQA DiamondGranite 4.0 1B by 4.423.728.1+4.4
- Humanity's Last Examtie5.34.8tie
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationGemma 3 1B Instruct by 11.217.76.5+11.2
- AA-Omniscience · AccuracyGranite 4.0 1B by 2.43.86.2+2.4
Other results18 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- BFCL v3Granite 4.0 1B by 35.816.652.4+35.8
- J-BFCLv3Granite 4.0 1B by 33.317.350.6+33.3
- JHumanEval+Granite 4.0 1B by 26.22551.2+26.2
- MMLU-ProGranite 4.0 1B by 18.814.733.5+18.8
- Domain AvgGranite 4.0 1B by 14.623.938.5+14.6
- τ²-Bench Telecom (AA run)Granite 4.0 1B by 12.310.522.8+12.3
- GSM8KGranite 4.0 1B by 10.662.873.4+10.6
- J-MATH500Granite 4.0 1B by 9.815.625.4+9.8
- Avg.Granite 4.0 1B by 9.524.634.1+9.5
- J-GSM8KGranite 4.0 1B by 9.233.642.8+9.2
- AA-OmniscienceGemma 3 1B Instruct by 6.2-75.5-81.6+6.2
- AA Agentic IndexGranite 4.0 1B by 4.13.57.6+4.1
- JMMLU-ProXGranite 4.0 1B by 1.214.115.3+1.2
- IFEvaltie80.279.6tie
- JMMLUtie34.533.9tie
- MATH-500 (EM)tie45.244.8tie
- AA Intelligencetie4.85tie
- JGPQAtie24.224.4tie
Questions people ask
Which is better, Gemma 3 1B Instruct or Granite 4.0 1B?
Granite 4.0 1B wins two of the five areas where both have results: reasoning and long documents. Gemma 3 1B Instruct wins none. They are level on coding, facts and following instructions.
Which is better for coding?
Neither. All 1 coding tests both models report are ties.
How do you compare the two?
We use the 26 benchmark tests both models have published scores on. The verdict counts the 8 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 18 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.