The two are evenly matched.LFM2.5-VL-3B is better at images and charts; Qwen3.5 4B at following instructions.
Scores updated · 29 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Following instructions rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where LFM2.5-VL-3B pulls ahead
Every test, side by side
All 29 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsLFM2.5-VL-3B
Following instructionsQwen3.5 4B
- IFBenchQwen3.5 4B by 26.225.852+26.2
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- BFCLv4Qwen3.5 4B by 18.132.550.6+18.1
- OCRBench v2_enQwen3.5 4B by 11.247.558.7+11.2
- InfographicVQA (val)Qwen3.5 4B by 10.170.280.3+10.1
- MMEQwen3.5 4B by 6.473.179.5+6.4
- RealWorldQALFM2.5-VL-3B by 673.167.1+6
- SimpleVQAQwen3.5 4B by 5.335.440.7+5.3
- HallusionBenchQwen3.5 4B by 4.547.251.7+4.5
- ScreenSpot-v2 WebLFM2.5-VL-3B by 4.482.277.8+4.4
- IFEvalQwen3.5 4B by 3.982.386.2+3.9
- DocVQA-valQwen3.5 4B by 3.791.194.8+3.7
- TextVQA-valLFM2.5-VL-3B by 3.184.381.2+3.1
- ChartQA TestQwen3.5 4B by 2.981.384.2+2.9
- POPELFM2.5-VL-3B by 2.788.786+2.7
- MMBenchLFM2.5-VL-3B by 2.68178.4+2.6
- MM-IFEvalQwen3.5 4B by 2.560.663.1+2.5
- Multilingual MMBenchLFM2.5-VL-3B by 2.579.577+2.5
- ScreenSpot-v2 DesktopLFM2.5-VL-3B by 2.478.776.3+2.4
- MMMU (val) (Pass@1)Qwen3.5 4B by 1.948.450.3+1.9
- SEED-Bench (image)LFM2.5-VL-3B by 1.677.776.1+1.6
- OCRBench v1Qwen3.5 4B by 1.484.285.6+1.4
- MMMBLFM2.5-VL-3B by 18382+1
- Multilingual MMMBLFM2.5-VL-3B by 18382+1
- ScreenSpot-v2 Mobiletie81.281.4tie
Questions people ask
Which is better, LFM2.5-VL-3B or Qwen3.5 4B?
LFM2.5-VL-3B and Qwen3.5 4B each win one of the two areas where both have results. LFM2.5-VL-3B wins images and charts; Qwen3.5 4B wins following instructions.
How do you compare the two?
We use the 29 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.