Granite 4.1 30B vs Granite 4.1 8B
Wins 4 of 6 areas
Coding · Reasoning · Long documents · Following instructions
Wins 1 of 6 areas
Agents
Granite 4.1 30B is the stronger all-rounder.Granite 4.1 8B is better at agents.
Scores updated · 44 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20Granite 4.1 30B2 of 3 tests · 1 tie
- ReasoningHard problems that need careful thinking10Granite 4.1 30B1 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts10Granite 4.1 30B1 of 1 test
- Following instructionsDoing exactly what it is asked10Granite 4.1 30B1 of 1 test
- AgentsCarrying out multi-step tasks on its own01Granite 4.1 8B1 of 2 tests · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Granite 4.1 30B pulls ahead
- Reasons across sets of long documentsAA-LCR+10.4points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+5.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4.8points ahead
Where Granite 4.1 8B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+7.9points ahead
Every test, side by side
All 44 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGranite 4.1 30B
- SciCodeGranite 4.1 30B by 425.821.8+4
- Terminal-Bench HardGranite 4.1 30B by 2.32.30+2.3
- Terminal-Bench 2.1tie2.63.4tie
AgentsGranite 4.1 8B
- GDPValGranite 4.1 8B by 2.202.2+2.2
- τ-Bench V3 · Bankingtie3.93.1tie
ReasoningGranite 4.1 30B
- GPQA DiamondGranite 4.1 30B by 4.848.143.3+4.8
- Humanity's Last Examtie4.13.8tie
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationGranite 4.1 8B by 7.9512.9+7.9
- AA-Omniscience · AccuracyGranite 4.1 30B by 1.713.912.3+1.7
Following instructionsGranite 4.1 30B
- IFBenchGranite 4.1 30B by 5.844.438.6+5.8
Other results32 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- τ²-Bench Telecom (AA run)Granite 4.1 30B by 14.342.127.8+14.3
- MGSMGranite 4.1 8B by 11.271.182.3+11.2
- MMMLUGranite 4.1 30B by 8.973.764.8+8.9
- IncludeGranite 4.1 30B by 8.467.358.9+8.4
- CRUXEval-O (output prediction)Granite 4.1 30B by 8.255.847.6+8.2
- MMLU-ProGranite 4.1 30B by 8.164.156+8.1
- GSM SymbolicGranite 4.1 8B by 875.783.7+8
- MMLUGranite 4.1 30B by 6.480.273.8+6.4
- AlpacaEval-2.0Granite 4.1 30B by 6.156.250.1+6.1
- HumanEval+Granite 4.1 30B by 5.585.479.9+5.5
- AGI EVALGranite 4.1 30B by 5.477.872.4+5.4
- BFCL v3Granite 4.1 30B by 5.473.768.3+5.4
- AttaQGranite 4.1 30B by 4.685.881.2+4.6
- BigCodeBenchGranite 4.1 30B by 3.838.835+3.8
- AA-OmniscienceGranite 4.1 8B by 3.7-67.8-64.1+3.7
- BBHGranite 4.1 30B by 3.283.780.5+3.2
- HumanEvalGranite 4.1 30B by 388.485.4+3
- IFEval (avg)Granite 4.1 30B by 2.689.787.1+2.6
- Tulu3 Safety Eval AvgGranite 4.1 30B by 2.678.275.6+2.6
- Eval+ AvgGranite 4.1 30B by 2.582.780.2+2.5
- Arena HardGranite 4.1 30B by 27169+2
- MultiPL-EGranite 4.1 30B by 262.360.3+2
- SimpleQAGranite 4.1 30B by 26.84.8+2
- DeepMind MathGranite 4.1 30B by 1.881.980.1+1.8
- MBPPGranite 4.1 8B by 1.885.587.3+1.8
- GSM8KGranite 4.1 30B by 1.794.292.5+1.7
- Minerva MathGranite 4.1 30B by 1.281.380.1+1.2
- Artificial Analysis Coding Indextie10.49.5tie
- AA Intelligencetie7.46.6tie
- SALAD-Benchtie96.495.8tie
- MBPP+ (EvalPlus-augmented)tie73.573.8tie
- MTBench Avgtie8.68.6tie
Questions people ask
Which is better, Granite 4.1 30B or Granite 4.1 8B?
Granite 4.1 30B wins four of the six areas where both have results: coding, reasoning, long documents and following instructions. Granite 4.1 8B wins agents. They are level on facts.
Which is better for coding?
Granite 4.1 30B. It wins 2 of the 3 coding tests both models report; Granite 4.1 8B wins none, and 1 is a tie.
How do you compare the two?
We use the 44 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 32 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.