The two are evenly matched.LFM2.5-VL-3B is better at images and charts; Qwen3.5 2B at following instructions.
Scores updated · 29 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Following instructions rests on a single test.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where LFM2.5-VL-3B pulls ahead
Every test, side by side
All 29 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsLFM2.5-VL-3B
Following instructionsQwen3.5 2B
- IFBenchQwen3.5 2B by 5.725.831.5+5.7
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- HallusionBenchQwen3.5 2B by 18.347.265.5+18.3
- ScreenSpot-v2 WebLFM2.5-VL-3B by 16.382.265.9+16.3
- ScreenSpot-v2 DesktopLFM2.5-VL-3B by 14.978.763.8+14.9
- ScreenSpot-v2 MobileLFM2.5-VL-3B by 11.581.269.7+11.5
- BFCLv4Qwen3.5 2B by 11.132.543.6+11.1
- MMBenchLFM2.5-VL-3B by 11.18169.9+11.1
- Multilingual MMBenchLFM2.5-VL-3B by 9.679.569.9+9.6
- IFEvalLFM2.5-VL-3B by 8.782.373.6+8.7
- MMMBLFM2.5-VL-3B by 8.58374.5+8.5
- RealWorldQALFM2.5-VL-3B by 873.165.1+8
- Multilingual MMMBLFM2.5-VL-3B by 7.18375.9+7.1
- TextVQA-valLFM2.5-VL-3B by 784.377.3+7
- MM-IFEvalLFM2.5-VL-3B by 5.260.655.4+5.2
- MMMU (val) (Pass@1)LFM2.5-VL-3B by 4.348.444.1+4.3
- InfographicVQA (val)Qwen3.5 2B by 3.370.273.5+3.3
- MMEQwen3.5 2B by 3.173.176.2+3.1
- ChartQA TestLFM2.5-VL-3B by 2.981.378.4+2.9
- SEED-Bench (image)LFM2.5-VL-3B by 1.977.775.8+1.9
- DocVQA-valQwen3.5 2B by 1.591.192.6+1.5
- OCRBench v2_entie47.548.1tie
- OCRBench v1tie84.284.4tie
- SimpleVQAtie35.435.2tie
- POPEtie88.788.6tie
Questions people ask
Which is better, LFM2.5-VL-3B or Qwen3.5 2B?
LFM2.5-VL-3B and Qwen3.5 2B each win one of the two areas where both have results. LFM2.5-VL-3B wins images and charts; Qwen3.5 2B wins following instructions.
How do you compare the two?
We use the 29 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.