Gemma 3 4B vs Qwen3 1.7B
Wins 2 of 6 areas
Long documents · Following instructions
Wins 2 of 6 areas
Reasoning · Facts
The two are evenly matched.Gemma 3 4B is better at long documents and following instructions; Qwen3 1.7B at facts and reasoning.
Scores updated · 25 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- FactsGetting facts right instead of making them up02Qwen3 1.7B2 of 2 tests
- ReasoningHard problems that need careful thinking01Qwen3 1.7B1 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts10Gemma 3 4B1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemma 3 4B1 of 1 test
- CodingWriting and fixing software11Even1 each · 1 tie
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3 1.7B pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+6.5points ahead
- Recent programming contest problemsLiveCodeBench v6+5.2points ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+3.4points ahead
Every test, side by side
All 25 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- LiveCodeBench v6Qwen3 1.7B by 5.218.924.1+5.2
- SciCodeGemma 3 4B by 37.34.3+3
- Terminal-Bench Hardtie0.80tie
ReasoningQwen3 1.7B
- GPQA DiamondQwen3 1.7B by 6.529.135.6+6.5
- Humanity's Last Examtie5.34.6tie
- CritPttie00tie
FactsQwen3 1.7B
- AA-Omniscience · Non-hallucinationQwen3 1.7B by 3.41.85.2+3.4
- AA-Omniscience · AccuracyQwen3 1.7B by 1.27.78.9+1.2
Following instructionsGemma 3 4B
- IFBenchGemma 3 4B by 1.428.326.9+1.4
Other results14 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GSM8KGemma 3 4B by 37.889.251.4+37.8
- Creative Writing v3Gemma 3 4B by 3768.631.6+37
- τ²-Bench Telecom (AA run)Qwen3 1.7B by 21526+21
- MGSMGemma 3 4B by 20.787.366.6+20.7
- IFEvalGemma 3 4B by 16.290.274+16.2
- MMLU-ProQwen3 1.7B by 13.443.657+13.4
- MATH-500 (EM)Qwen3 1.7B by 8.773.281.9+8.7
- LCB v5Qwen3 1.7B by 7.419.126.5+7.4
- AA Agentic IndexQwen3 1.7B by 71.78.7+7
- AA-OmniscienceQwen3 1.7B by 5.4-82.9-77.5+5.4
- MMMLUGemma 3 4B by 3.650.146.5+3.6
- HumanEval+Gemma 3 4B by 1.862.861+1.8
- MMLUtie58.459.1tie
- AA Intelligencetie4.85.2tie
Questions people ask
Which is better, Gemma 3 4B or Qwen3 1.7B?
Gemma 3 4B and Qwen3 1.7B each win two of the six areas where both have results. Gemma 3 4B wins long documents and following instructions; Qwen3 1.7B wins reasoning and facts. They are level on coding and agents.
Which is better for coding?
Neither. They win 1 coding test each of the 3 both models report, and 1 is a tie.
How do you compare the two?
We use the 25 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 14 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.