LFM2-VL-3B vs Qwen3.5 4B
Wins 0 of 2 areas
—
Wins 2 of 2 areas
Images and charts · Following instructions
Qwen3.5 4B is the stronger all-rounder.
Scores updated · 29 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Following instructions rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3.5 4B pulls ahead
Where LFM2-VL-3B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 29 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsQwen3.5 4B
Following instructionsQwen3.5 4B
- IFBenchQwen3.5 4B by 31.220.852+31.2
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ScreenSpot-v2 WebQwen3.5 4B by 75.32.577.8+75.3
- ScreenSpot-v2 MobileQwen3.5 4B by 73.87.681.4+73.8
- ScreenSpot-v2 DesktopQwen3.5 4B by 70.3676.3+70.3
- BFCLv4Qwen3.5 4B by 30.120.550.6+30.1
- OCRBench v2_enQwen3.5 4B by 14.843.958.7+14.8
- IFEvalQwen3.5 4B by 13.372.986.2+13.3
- InfographicVQA (val)Qwen3.5 4B by 12.567.880.3+12.5
- MM-IFEvalQwen3.5 4B by 11.751.463.1+11.7
- SimpleVQAQwen3.5 4B by 7.73340.7+7.7
- MMEQwen3.5 4B by 6.57379.5+6.5
- HallusionBenchQwen3.5 4B by 5.346.451.7+5.3
- DocVQA-valQwen3.5 4B by 589.894.8+5
- MMMU (val) (Pass@1)Qwen3.5 4B by 4.745.650.3+4.7
- RealWorldQALFM2-VL-3B by 471.167.1+4
- OCRBench v1Qwen3.5 4B by 3.981.785.6+3.9
- ChartQA TestQwen3.5 4B by 3.880.484.2+3.8
- POPELFM2-VL-3B by 3.289.286+3.2
- TextVQA-valLFM2-VL-3B by 1.88381.2+1.8
- MMBenchLFM2-VL-3B by 1.68078.4+1.6
- Multilingual MMBenchtie76.377tie
- SEED-Bench (image)tie76.676.1tie
- MMMBtie81.982tie
- Multilingual MMMBtie81.982tie
Questions people ask
Which is better, LFM2-VL-3B or Qwen3.5 4B?
Qwen3.5 4B wins all two areas where both have results: images and charts and following instructions. LFM2-VL-3B wins none.
How do you compare the two?
We use the 29 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.