Granite 4.0 350M vs Qwen3.5 4B
Wins 0 of 5 areas
—
Wins 5 of 5 areas
Coding · Reasoning · Facts · Long documents · Following instructions
Qwen3.5 4B is the stronger all-rounder.
Scores updated · 18 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking02Qwen3.5 4B2 of 3 tests · 1 tie
- CodingWriting and fixing software02Qwen3.5 4B2 of 2 tests
- FactsGetting facts right instead of making them up02Qwen3.5 4B2 of 2 tests
- Long documentsFinding answers in very long texts01Qwen3.5 4B1 of 1 test
- Following instructionsDoing exactly what it is asked01Qwen3.5 4B1 of 1 test
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3.5 4B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+51.4points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+34.4points ahead
- Code for real scientific research problemsSciCode+14.4points ahead
Where Granite 4.0 350M pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 18 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3.5 4B
- Terminal-Bench HardQwen3.5 4B by 18.2018.2+18.2
- SciCodeQwen3.5 4B by 14.41.716.1+14.4
ReasoningQwen3.5 4B
- GPQA DiamondQwen3.5 4B by 51.425.777.1+51.4
- Humanity's Last ExamQwen3.5 4B by 3.56.49.9+3.5
- CritPttie00tie
FactsQwen3.5 4B
- AA-Omniscience · AccuracyQwen3.5 4B by 11.33.815.1+11.3
- AA-Omniscience · Non-hallucinationQwen3.5 4B by 1.511.913.4+1.5
Following instructionsQwen3.5 4B
- IFBenchQwen3.5 4B by 34.417.652+34.4
Other results9 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- τ²-Bench Telecom (AA run)Qwen3.5 4B by 77.514.692.1+77.5
- MMLU-ProQwen3.5 4B by 66.312.879.1+66.3
- τ²-Bench (Retail)Qwen3.5 4B by 65.86.171.9+65.8
- BFCLv4Qwen3.5 4B by 36.613.750.3+36.6
- IFEvalQwen3.5 4B by 32.753.586.2+32.7
- BFCL v3Qwen3.5 4B by 31.539.671.1+31.5
- AA Agentic IndexQwen3.5 4B by 27.64.932.5+27.6
- AA-OmniscienceQwen3.5 4B by 22.6-80.9-58.4+22.6
- AA IntelligenceQwen3.5 4B by 8.34.813.1+8.3
Questions people ask
Which is better, Granite 4.0 350M or Qwen3.5 4B?
Qwen3.5 4B wins all five areas where both have results: coding, reasoning, facts, long documents and following instructions. Granite 4.0 350M wins none.
Which is better for coding?
Qwen3.5 4B. It wins 2 of the 2 coding tests both models report; Granite 4.0 350M wins none.
How do you compare the two?
We use the 18 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 9 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.