Gemma 4 E2B vs Phi 4 Multimodal Instruct
Wins 2 of 3 areas
Coding · Reasoning
Wins 1 of 3 areas
Images and charts
Gemma 4 E2B is the stronger all-rounder.Phi 4 Multimodal Instruct is better at images and charts.
Scores updated · 15 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding rests on a single test.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemma 4 E2B pulls ahead
- Harder college exam questions with imagesMMMU-Pro+30.1points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+11.8points ahead
- Code for real scientific research problemsSciCode+9.9points ahead
Every test, side by side
All 15 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningGemma 4 E2B
- GPQA DiamondGemma 4 E2B by 11.843.331.5+11.8
- Humanity's Last Examtie4.85tie
Images and chartsPhi 4 Multimodal Instruct
Other results6 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ChartQAPhi 4 Multimodal Instruct by 37.943.581.4+37.9
- TextVQA-valGemma 4 E2B by 22.662.539.9+22.6
- MMBenchPhi 4 Multimodal Instruct by 22.564.286.7+22.5
- DocVQA-valPhi 4 Multimodal Instruct by 7.185.792.8+7.1
- AA IntelligenceGemma 4 E2B by 27.85.8+2
- POPEPhi 4 Multimodal Instruct by 1.68485.6+1.6
Questions people ask
Which is better, Gemma 4 E2B or Phi 4 Multimodal Instruct?
Gemma 4 E2B wins two of the three areas where both have results: coding and reasoning. Phi 4 Multimodal Instruct wins images and charts.
Which is better for coding?
Gemma 4 E2B. It wins the one coding test both models report.
How do you compare the two?
We use the 15 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 6 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.