Gemma 3 1B Instruct vs Qwen3 1.7B
Wins 0 of 5 areas
—
Wins 3 of 5 areas
Coding · Reasoning · Following instructions
Qwen3 1.7B is the stronger all-rounder.
Scores updated · 23 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking01Qwen3 1.7B1 of 3 tests · 2 ties
- CodingWriting and fixing software01Qwen3 1.7B1 of 2 tests · 1 tie
- Following instructionsDoing exactly what it is asked01Qwen3 1.7B1 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
- Long documentsFinding answers in very long texts00Even0 each · 1 tie
Long documents and following instructions rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3 1.7B pulls ahead
- Recent programming contest problemsLiveCodeBench v6+19.8points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+11.9points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+7points ahead
Where Gemma 3 1B Instruct pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+12.5points ahead
Every test, side by side
All 23 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3 1.7B
- LiveCodeBench v6Qwen3 1.7B by 19.84.324.1+19.8
- Terminal-Bench Hardtie00tie
ReasoningQwen3 1.7B
- GPQA DiamondQwen3 1.7B by 11.923.735.6+11.9
- Humanity's Last Examtie5.34.6tie
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationGemma 3 1B Instruct by 12.517.75.2+12.5
- AA-Omniscience · AccuracyQwen3 1.7B by 5.13.88.9+5.1
Other results14 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MMLU-ProQwen3 1.7B by 42.314.757+42.3
- BFCL v3Qwen3 1.7B by 38.816.655.4+38.8
- MATH-500 (EM)Qwen3 1.7B by 36.745.281.9+36.7
- HumanEval+Qwen3 1.7B by 23.837.261+23.8
- MGSMQwen3 1.7B by 2343.666.6+23
- LCB v5Qwen3 1.7B by 22.14.426.5+22.1
- MMLUQwen3 1.7B by 1940.159.1+19
- τ²-Bench Telecom (AA run)Qwen3 1.7B by 15.510.526+15.5
- MMMLUQwen3 1.7B by 12.134.446.5+12.1
- GSM8KGemma 3 1B Instruct by 11.462.851.4+11.4
- IFEvalGemma 3 1B Instruct by 6.280.274+6.2
- AA Agentic IndexQwen3 1.7B by 5.23.58.7+5.2
- AA-OmniscienceGemma 3 1B Instruct by 2.1-75.5-77.5+2.1
- AA Intelligencetie4.85.2tie
Questions people ask
Which is better, Gemma 3 1B Instruct or Qwen3 1.7B?
Qwen3 1.7B wins three of the five areas where both have results: coding, reasoning and following instructions. Gemma 3 1B Instruct wins none. They are level on facts and long documents.
Which is better for coding?
Qwen3 1.7B. It wins 1 of the 2 coding tests both models report; Gemma 3 1B Instruct wins none, and 1 is a tie.
How do you compare the two?
We use the 23 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 14 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.