LFM2.5-VL-1.6B vs Qwen3 VL 4B Instruct
Wins 1 of 7 areas
Following instructions
Wins 3 of 7 areas
Coding · Images and charts · Long documents
Qwen3 VL 4B Instruct wins more areas, narrowly.LFM2.5-VL-1.6B is better at following instructions.
Scores updated · 23 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- Images and chartsUnderstanding pictures, charts and video04Qwen3 VL 4B Instruct4 of 4 tests
- CodingWriting and fixing software01Qwen3 VL 4B Instruct1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Qwen3 VL 4B Instruct1 of 1 test
- Following instructionsDoing exactly what it is asked10LFM2.5-VL-1.6B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- ReasoningHard problems that need careful thinking11Even1 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
Agents, long documents and following instructions rest on a single test each.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Qwen3 VL 4B Instruct pulls ahead
Where LFM2.5-VL-1.6B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+1.7points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+1.5points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+1.3points ahead
Every test, side by side
All 23 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3 VL 4B Instruct
- SciCodeQwen3 VL 4B Instruct by 10.7313.7+10.7
- Terminal-Bench Hardtie00tie
ReasoningEven
- GPQA DiamondQwen3 VL 4B Instruct by 8.228.937.1+8.2
- Humanity's Last ExamLFM2.5-VL-1.6B by 1.55.13.6+1.5
- CritPttie00tie
FactsEven
- AA-Omniscience · AccuracyQwen3 VL 4B Instruct by 5.15.810.9+5.1
- AA-Omniscience · Non-hallucinationLFM2.5-VL-1.6B by 1.74.32.6+1.7
Images and chartsQwen3 VL 4B Instruct
Following instructionsLFM2.5-VL-1.6B
- IFBenchLFM2.5-VL-1.6B by 1.333.131.8+1.3
Other results9 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MMMU (val) (Pass@1)Qwen3 VL 4B Instruct by 26.840.667.4+26.8
- VLM Judge ScoreQwen3 VL 4B Instruct by 266692+26
- OCRBench v2_enQwen3 VL 4B Instruct by 22.241.563.7+22.2
- τ²-Bench Telecom (AA run)Qwen3 VL 4B Instruct by 14.98.523.4+14.9
- AA-OmniscienceQwen3 VL 4B Instruct by 8.4-84.4-76+8.4
- RealWorldQAQwen3 VL 4B Instruct by 6.164.870.9+6.1
- AA Agentic IndexQwen3 VL 4B Instruct by 52.87.8+5
- HallusionBenchLFM2.5-VL-1.6B by 2.560.157.6+2.5
- AA Intelligencetie4.85.7tie
Questions people ask
Which is better, LFM2.5-VL-1.6B or Qwen3 VL 4B Instruct?
Qwen3 VL 4B Instruct wins three of the seven areas where both have results: coding, images and charts and long documents. LFM2.5-VL-1.6B wins following instructions. They are level on agents, reasoning and facts.
Which is better for coding?
Qwen3 VL 4B Instruct. It wins 1 of the 2 coding tests both models report; LFM2.5-VL-1.6B wins none, and 1 is a tie.
How do you compare the two?
We use the 23 benchmark tests both models have published scores on. The verdict counts the 14 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 9 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.