Mage-VL-4B is the stronger all-rounder.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the one area where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Mage-VL-4B pulls ahead
Where Phi 4 Multimodal Instruct pulls ahead
- Reads text in imagesOCRBench+2.6points ahead
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsMage-VL-4B
Other results20 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Timelens-ActivityNetMage-VL-4B by 43.445.42+43.4
- MultiDocVQA-valMage-VL-4B by 40.787.546.8+40.7
- VSI-BenchMage-VL-4B by 40.264.324.1+40.2
- TextVQA-valMage-VL-4B by 37.477.339.9+37.4
- MME-RealWorldMage-VL-4B by 3466.532.5+34
- CV-BenchMage-VL-4B by 30.787.857.1+30.7
- MME-PerceptionMage-VL-4B by 300 rating points17101410+300 rating
- NextQAMage-VL-4B by 2983.154.1+29
- MLVU-devMage-VL-4B by 24.568.744.2+24.5
- Ref-DAVIS17Mage-VL-4B by 22.725.83.1+22.7
- LongVideoBenchMage-VL-4B by 20.261.341.1+20.2
- MV-BenchMage-VL-4B by 20.265.144.9+20.2
- MMBench-EN-devMage-VL-4B by 18.28465.8+18.2
- LVBenchMage-VL-4B by 16.541.825.3+16.5
- SATMage-VL-4B by 1267.355.3+12
- InfoVQA (val)Mage-VL-4B by 8.580.371.8+8.5
- MMBench-CN-devMage-VL-4B by 6.88275.2+6.8
- ChartQAMage-VL-4B by 3.584.981.4+3.5
- DocVQA-valMage-VL-4B by 2.395.192.8+2.3
- AI2D w/ MaskMage-VL-4B by 1.483.281.8+1.4
Questions people ask
Which is better, Mage-VL-4B or Phi 4 Multimodal Instruct?
Mage-VL-4B wins the one area where both have results: images and charts. Phi 4 Multimodal Instruct wins none.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 4 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 20 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.