Granite 4.1 30B vs Qwen2.5 Instruct 72B
Wins 2 of 5 areas
Long documents · Following instructions
Wins 3 of 5 areas
Coding · Reasoning · Facts
Qwen2.5 Instruct 72B is the stronger all-rounder.Granite 4.1 30B is better at long documents.
Scores updated · 27 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- FactsGetting facts right instead of making them up02Qwen2.5 Instruct 72B2 of 2 tests
- ReasoningHard problems that need careful thinking01Qwen2.5 Instruct 72B1 of 3 tests · 2 ties
- CodingWriting and fixing software01Qwen2.5 Instruct 72B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts10Granite 4.1 30B1 of 1 test
- Following instructionsDoing exactly what it is asked10Granite 4.1 30B1 of 1 test
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen2.5 Instruct 72B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+9.5points ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+3.5points ahead
- Hard command-line tasks in a real terminalTerminal-Bench Hard+2.2points ahead
Every test, side by side
All 27 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen2.5 Instruct 72B
- Terminal-Bench HardQwen2.5 Instruct 72B by 2.22.34.5+2.2
- SciCodetie25.826.7tie
ReasoningQwen2.5 Instruct 72B
- GPQA DiamondQwen2.5 Instruct 72B by 148.149.1+1
- Humanity's Last Examtie4.13.6tie
- CritPttie00tie
FactsQwen2.5 Instruct 72B
- AA-Omniscience · Non-hallucinationQwen2.5 Instruct 72B by 9.5514.5+9.5
- AA-Omniscience · AccuracyQwen2.5 Instruct 72B by 3.513.917.5+3.5
Following instructionsGranite 4.1 30B
- IFBenchGranite 4.1 30B by 7.544.436.9+7.5
Other results18 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceQwen2.5 Instruct 72B by 14.7-67.8-53.1+14.7
- MultiPL-EQwen2.5 Instruct 72B by 12.862.375.1+12.8
- Arena HardQwen2.5 Instruct 72B by 10.27181.2+10.2
- τ²-Bench Telecom (AA run)Granite 4.1 30B by 7.642.134.5+7.6
- MMLU-ProQwen2.5 Instruct 72B by 764.171.1+7
- MGSMQwen2.5 Instruct 72B by 5.171.176.2+5.1
- CRUXEval-O (output prediction)Qwen2.5 Instruct 72B by 4.155.859.9+4.1
- BBHGranite 4.1 30B by 3.983.779.8+3.9
- MBPP+ (EvalPlus-augmented)Qwen2.5 Instruct 72B by 3.573.577+3.5
- MMLUGranite 4.1 30B by 3.280.277+3.2
- MBPPQwen2.5 Instruct 72B by 2.785.588.2+2.7
- IFEval (avg)Granite 4.1 30B by 2.589.787.2+2.5
- SimpleQAQwen2.5 Instruct 72B by 2.36.89.1+2.3
- HumanEvalGranite 4.1 30B by 1.888.486.6+1.8
- GSM8KQwen2.5 Instruct 72B by 1.694.295.8+1.6
- Artificial Analysis Coding IndexQwen2.5 Instruct 72B by 1.510.411.9+1.5
- MMMLUQwen2.5 Instruct 72B by 1.173.774.8+1.1
- AA Intelligencetie7.47.7tie
Questions people ask
Which is better, Granite 4.1 30B or Qwen2.5 Instruct 72B?
Qwen2.5 Instruct 72B wins three of the five areas where both have results: coding, reasoning and facts. Granite 4.1 30B wins long documents and following instructions.
Which is better for coding?
Qwen2.5 Instruct 72B. It wins 1 of the 2 coding tests both models report; Granite 4.1 30B wins none, and 1 is a tie.
How do you compare the two?
We use the 27 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 18 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.