Granite 4.0 350M vs Qwen3.5 0.8B
Wins 2 of 6 areas
Coding · Reasoning
Wins 3 of 6 areas
Facts · Long documents · Following instructions
Qwen3.5 0.8B wins more areas, narrowly.Granite 4.0 350M is better at reasoning.
Scores updated · 20 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- FactsGetting facts right instead of making them up01Qwen3.5 0.8B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Qwen3.5 0.8B1 of 1 test
- Following instructionsDoing exactly what it is asked01Qwen3.5 0.8B1 of 1 test
- ReasoningHard problems that need careful thinking20Granite 4.0 350M2 of 3 tests · 1 tie
- CodingWriting and fixing software10Granite 4.0 350M1 of 2 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
Agents, long documents and following instructions rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3.5 0.8B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+26.8points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+3.9points ahead
Where Granite 4.0 350M pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+14.6points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+5.3points ahead
Every test, side by side
All 20 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGranite 4.0 350M
- SciCodeGranite 4.0 350M by 1.71.70+1.7
- Terminal-Bench Hardtie00tie
ReasoningGranite 4.0 350M
- GPQA DiamondGranite 4.0 350M by 14.625.711.1+14.6
- Humanity's Last ExamGranite 4.0 350M by 5.36.41.1+5.3
- CritPttie00tie
FactsQwen3.5 0.8B
- AA-Omniscience · Non-hallucinationQwen3.5 0.8B by 26.811.938.7+26.8
- AA-Omniscience · Accuracytie3.84.2tie
Following instructionsQwen3.5 0.8B
- IFBenchQwen3.5 0.8B by 3.917.621.5+3.9
Other results10 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- τ²-Bench Telecom (AA run)Qwen3.5 0.8B by 33.114.647.7+33.1
- MMLU-ProQwen3.5 0.8B by 29.512.842.3+29.5
- AA-OmniscienceQwen3.5 0.8B by 26.4-80.9-54.5+26.4
- IFEvalQwen3.5 0.8B by 17.753.571.2+17.7
- BFCLv4Qwen3.5 0.8B by 11.613.725.3+11.6
- τ²-BenchQwen3.5 0.8B by 11.42.914.3+11.4
- AA IntelligenceQwen3.5 0.8B by 1.34.86.1+1.3
- τ²-Bench (Retail)tie6.17tie
- AA Agentic Indextie4.95.7tie
- BFCL v3tie39.639.6tie
Questions people ask
Which is better, Granite 4.0 350M or Qwen3.5 0.8B?
Qwen3.5 0.8B wins three of the six areas where both have results: facts, long documents and following instructions. Granite 4.0 350M wins coding and reasoning. They are level on agents.
Which is better for coding?
Granite 4.0 350M. It wins 1 of the 2 coding tests both models report; Qwen3.5 0.8B wins none, and 1 is a tie.
How do you compare the two?
We use the 20 benchmark tests both models have published scores on. The verdict counts the 10 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 10 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.