The two are evenly matched.LFM2-VL-3B is better at images and charts; Qwen3.5 2B at following instructions.
Scores updated · 29 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Following instructions rests on a single test.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where LFM2-VL-3B pulls ahead
Every test, side by side
All 29 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsLFM2-VL-3B
Following instructionsQwen3.5 2B
- IFBenchQwen3.5 2B by 10.720.831.5+10.7
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ScreenSpot-v2 WebQwen3.5 2B by 63.42.565.9+63.4
- ScreenSpot-v2 MobileQwen3.5 2B by 62.17.669.7+62.1
- ScreenSpot-v2 DesktopQwen3.5 2B by 57.8663.8+57.8
- HallusionBenchQwen3.5 2B by 19.146.465.5+19.1
- BFCLv4Qwen3.5 2B by 13.420.533.9+13.4
- MMMBLFM2-VL-3B by 7.481.974.5+7.4
- MMBenchLFM2-VL-3B by 6.98073.1+6.9
- Multilingual MMBenchLFM2-VL-3B by 6.476.369.9+6.4
- Multilingual MMMBLFM2-VL-3B by 681.975.9+6
- RealWorldQALFM2-VL-3B by 671.165.1+6
- InfographicVQA (val)Qwen3.5 2B by 5.767.873.5+5.7
- TextVQA-valLFM2-VL-3B by 5.78377.3+5.7
- OCRBench v2_enQwen3.5 2B by 4.243.948.1+4.2
- MM-IFEvalQwen3.5 2B by 451.455.4+4
- MMEQwen3.5 2B by 3.27376.2+3.2
- DocVQA-valQwen3.5 2B by 2.889.892.6+2.8
- OCRBench v1Qwen3.5 2B by 2.781.784.4+2.7
- SimpleVQAQwen3.5 2B by 2.23335.2+2.2
- ChartQA TestLFM2-VL-3B by 280.478.4+2
- MMMU (val) (Pass@1)LFM2-VL-3B by 1.545.644.1+1.5
- SEED-Bench (image)tie76.675.8tie
- IFEvaltie72.973.6tie
- POPEtie89.288.6tie
Questions people ask
Which is better, LFM2-VL-3B or Qwen3.5 2B?
LFM2-VL-3B and Qwen3.5 2B each win one of the two areas where both have results. LFM2-VL-3B wins images and charts; Qwen3.5 2B wins following instructions.
How do you compare the two?
We use the 29 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.