LFM2.5-VL-3B vs LFM2-VL-3B
Wins 2 of 2 areas
Images and charts · Following instructions
Wins 0 of 2 areas
—
LFM2.5-VL-3B is the stronger all-rounder.
Scores updated · 32 tests both models report · How we compare
Where each one wins
Tests won in each of the two areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Following instructions rests on a single test.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where LFM2.5-VL-3B pulls ahead
Where LFM2-VL-3B pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 32 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Images and chartsLFM2.5-VL-3B
Following instructionsLFM2.5-VL-3B
- IFBenchLFM2.5-VL-3B by 525.820.8+5
Other results25 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- ScreenSpot-v2 WebLFM2.5-VL-3B by 79.782.22.5+79.7
- ScreenSpot-v2 MobileLFM2.5-VL-3B by 73.681.27.6+73.6
- ScreenSpot-v2 DesktopLFM2.5-VL-3B by 72.778.76+72.7
- RefCOCO (Macro Prec@1)LFM2.5-VL-3B by 30.887.957.1+30.8
- BFCLv4LFM2.5-VL-3B by 1232.520.5+12
- IFEvalLFM2.5-VL-3B by 9.482.372.9+9.4
- MM-IFEvalLFM2.5-VL-3B by 9.260.651.4+9.2
- OCRBench v2_enLFM2.5-VL-3B by 3.647.543.9+3.6
- Multilingual MMBenchLFM2.5-VL-3B by 3.279.576.3+3.2
- MMMU (val) (Pass@1)LFM2.5-VL-3B by 2.848.445.6+2.8
- OCRBench v1LFM2.5-VL-3B by 2.584.281.7+2.5
- InfographicVQA (val)LFM2.5-VL-3B by 2.470.267.8+2.4
- SimpleVQALFM2.5-VL-3B by 2.435.433+2.4
- RealWorldQALFM2.5-VL-3B by 273.171.1+2
- DocVQA-valLFM2.5-VL-3B by 1.391.189.8+1.3
- TextVQA-valLFM2.5-VL-3B by 1.384.383+1.3
- MMMBLFM2.5-VL-3B by 1.18381.9+1.1
- Multilingual MMMBLFM2.5-VL-3B by 1.18381.9+1.1
- SEED-Bench (image)LFM2.5-VL-3B by 1.177.776.6+1.1
- MMBenchLFM2.5-VL-3B by 18180+1
- ChartQAtie81.380.4tie
- ChartQA Testtie81.380.4tie
- HallusionBenchtie47.246.4tie
- POPEtie88.789.2tie
- MMEtie73.173tie
Questions people ask
Which is better, LFM2.5-VL-3B or LFM2-VL-3B?
LFM2.5-VL-3B wins all two areas where both have results: images and charts and following instructions. LFM2-VL-3B wins none.
How do you compare the two?
We use the 32 benchmark tests both models have published scores on. The verdict counts the 7 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 25 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.