LFM2.5-VL-1.6B vs Qwen3.5 2B
Wins 1 of 7 areas
Following instructions
Wins 4 of 7 areas
Coding · Facts · Images and charts · Long documents
Qwen3.5 2B wins more areas, narrowly.LFM2.5-VL-1.6B is better at following instructions.
Scores updated · 33 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- Images and chartsUnderstanding pictures, charts and video03Qwen3.5 2B3 of 4 tests · 1 tie
- FactsGetting facts right instead of making them up02Qwen3.5 2B2 of 2 tests
- CodingWriting and fixing software01Qwen3.5 2B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Qwen3.5 2B1 of 1 test
- Following instructionsDoing exactly what it is asked10LFM2.5-VL-1.6B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- ReasoningHard problems that need careful thinking11Even1 each · 1 tie
Agents, long documents and following instructions rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3.5 2B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+21.5points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+16.7points ahead
- Harder college exam questions with imagesMMMU-Pro+16.5points ahead
Where LFM2.5-VL-1.6B pulls ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+2.5points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+1.6points ahead
Every test, side by side
All 33 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningEven
- GPQA DiamondQwen3.5 2B by 16.728.945.6+16.7
- Humanity's Last ExamLFM2.5-VL-1.6B by 2.55.12.6+2.5
- CritPttie00tie
FactsQwen3.5 2B
- AA-Omniscience · Non-hallucinationQwen3.5 2B by 21.54.325.8+21.5
- AA-Omniscience · AccuracyQwen3.5 2B by 2.35.88.1+2.3
Images and chartsQwen3.5 2B
Following instructionsLFM2.5-VL-1.6B
- IFBenchLFM2.5-VL-1.6B by 1.633.131.5+1.6
Other results19 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- τ²-Bench Telecom (AA run)Qwen3.5 2B by 60.58.569+60.5
- AA-OmniscienceQwen3.5 2B by 24.2-84.4-60.2+24.2
- VLM Judge ScoreQwen3.5 2B by 23.76689.7+23.7
- GQA TestDev_BalancedQwen3.5 2B by 14.439.553.9+14.4
- MMLU-Pro testQwen3.5 2B by 9.919.929.8+9.9
- MMMU DEV_VALQwen3.5 2B by 9.43847.4+9.4
- MMLU testQwen3.5 2B by 7.946.454.3+7.9
- OCRBench v2_enQwen3.5 2B by 6.641.548.1+6.6
- MMBench DEV_EN_V11Qwen3.5 2B by 6.469.676+6.4
- HallusionBenchQwen3.5 2B by 5.460.165.5+5.4
- DocVQA-valQwen3.5 2B by 4.987.792.6+4.9
- MTL MMBench_DEVQwen3.5 2B by 4.662.366.9+4.6
- ChartQA TestQwen3.5 2B by 4.573.978.4+4.5
- MMMU (val) (Pass@1)Qwen3.5 2B by 3.540.644.1+3.5
- MM-IFEvalQwen3.5 2B by 3.152.355.4+3.1
- MMMBQwen3.5 2B by 2.871.774.5+2.8
- AA IntelligenceQwen3.5 2B by 2.14.86.9+2.1
- RealWorldQAtie64.865.1tie
- MMBench_DEVtie0.60.7tie
Questions people ask
Which is better, LFM2.5-VL-1.6B or Qwen3.5 2B?
Qwen3.5 2B wins four of the seven areas where both have results: coding, facts, images and charts and long documents. LFM2.5-VL-1.6B wins following instructions. They are level on agents and reasoning.
Which is better for coding?
Qwen3.5 2B. It wins 1 of the 2 coding tests both models report; LFM2.5-VL-1.6B wins none, and 1 is a tie.
How do you compare the two?
We use the 33 benchmark tests both models have published scores on. The verdict counts the 14 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 19 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.