Gemma 3 4B vs Phi 4 Mini Instruct
Wins 1 of 6 areas
Following instructions
Wins 4 of 6 areas
Coding · Reasoning · Facts · Long documents
Phi 4 Mini Instruct is the stronger all-rounder.Gemma 3 4B is better at following instructions.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking01Phi 4 Mini Instruct1 of 3 tests · 2 ties
- FactsGetting facts right instead of making them up12Phi 4 Mini Instruct2 of 3 tests
- CodingWriting and fixing software01Phi 4 Mini Instruct1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Phi 4 Mini Instruct1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemma 3 4B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 2 ties
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Phi 4 Mini Instruct pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+20.2points ahead
- Reasons across sets of long documentsAA-LCR+8.6points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4points ahead
Where Gemma 3 4B pulls ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+7.2points ahead
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingPhi 4 Mini Instruct
- SciCodePhi 4 Mini Instruct by 3.57.310.8+3.5
- Terminal-Bench Hardtie0.80tie
ReasoningPhi 4 Mini Instruct
- GPQA DiamondPhi 4 Mini Instruct by 429.133.1+4
- Humanity's Last Examtie5.34.5tie
- CritPttie00tie
FactsPhi 4 Mini Instruct
- AA-Omniscience · Non-hallucinationPhi 4 Mini Instruct by 20.21.822+20.2
- Vectara HHEM hallucination ratelower is betterGemma 3 4B by 17.16.423.5+17.1
- AA-Omniscience · AccuracyPhi 4 Mini Instruct by 1.87.79.5+1.8
Long documentsPhi 4 Mini Instruct
- AA-LCRPhi 4 Mini Instruct by 8.66.715.3+8.6
Following instructionsGemma 3 4B
- IFBenchGemma 3 4B by 7.228.321.1+7.2
Other results12 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- vectara_avg_summary_lengthPhi 4 Mini Instruct by 342.8 rating points77.4420.2+342.8 rating
- vectara_answer_ratePhi 4 Mini Instruct by 25.267.392.5+25.2
- MGSMGemma 3 4B by 23.487.363.9+23.4
- AA-OmnisciencePhi 4 Mini Instruct by 21.8-82.9-61.1+21.8
- vectara_factual_consistencyGemma 3 4B by 17.193.676.5+17.1
- MMLU-ProPhi 4 Mini Instruct by 9.243.652.8+9.2
- τ²-Bench Telecom (AA run)Phi 4 Mini Instruct by 3.258.2+3.2
- BBHGemma 3 4B by 1.872.270.4+1.8
- AA IntelligencePhi 4 Mini Instruct by 1.54.86.3+1.5
- Artificial Analysis Coding IndexPhi 4 Mini Instruct by 1.12.73.8+1.1
- Arena HardPhi 4 Mini Instruct by 131.832.8+1
- GSM8Ktie89.288.6tie
Questions people ask
Which is better, Gemma 3 4B or Phi 4 Mini Instruct?
Phi 4 Mini Instruct wins four of the six areas where both have results: coding, reasoning, facts and long documents. Gemma 3 4B wins following instructions. They are level on agents.
Which is better for coding?
Phi 4 Mini Instruct. It wins 1 of the 2 coding tests both models report; Gemma 3 4B wins none, and 1 is a tie.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 12 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.