LFM2.5-VL-1.6B vs Ministral 3 3B
Wins 1 of 7 areas
Following instructions
Wins 4 of 7 areas
Coding · Reasoning · Facts · Long documents
Ministral 3 3B wins more areas, narrowly.LFM2.5-VL-1.6B is better at following instructions.
Scores updated · 27 tests both models report · How we compare
Where each one wins
Tests won in each of the seven areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- FactsGetting facts right instead of making them up02Ministral 3 3B2 of 2 tests
- ReasoningHard problems that need careful thinking01Ministral 3 3B1 of 3 tests · 2 ties
- CodingWriting and fixing software01Ministral 3 3B1 of 2 tests · 1 tie
- Long documentsFinding answers in very long texts01Ministral 3 3B1 of 1 test
- Following instructionsDoing exactly what it is asked10LFM2.5-VL-1.6B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- Images and chartsUnderstanding pictures, charts and video11Even1 each
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Ministral 3 3B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+15.5points ahead
- Code for real scientific research problemsSciCode+12.3points ahead
- Harder college exam questions with imagesMMMU-Pro+11.6points ahead
Every test, side by side
All 27 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingMinistral 3 3B
- SciCodeMinistral 3 3B by 12.3315.3+12.3
- Terminal-Bench Hardtie00tie
ReasoningMinistral 3 3B
- GPQA DiamondMinistral 3 3B by 6.928.935.8+6.9
- Humanity's Last Examtie5.15.4tie
- CritPttie00tie
FactsMinistral 3 3B
- AA-Omniscience · Non-hallucinationMinistral 3 3B by 15.54.319.8+15.5
- AA-Omniscience · AccuracyMinistral 3 3B by 3.25.89+3.2
Images and chartsEven
Following instructionsLFM2.5-VL-1.6B
- IFBenchLFM2.5-VL-1.6B by 6.333.126.8+6.3
Other results15 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MMLU-Pro testMinistral 3 3B by 27.619.947.5+27.6
- AA-OmniscienceMinistral 3 3B by 20.4-84.4-64+20.4
- MMLU testMinistral 3 3B by 19.646.466+19.6
- τ²-Bench Telecom (AA run)Ministral 3 3B by 16.48.524.9+16.4
- GQA TestDev_BalancedMinistral 3 3B by 14.939.554.4+14.9
- MMMU DEV_VALMinistral 3 3B by 12.83850.8+12.8
- ChartQA TestMinistral 3 3B by 5.273.979.1+5.2
- HallusionBenchMinistral 3 3B by 5.160.165.2+5.1
- MTL MMBench_DEVMinistral 3 3B by 5.162.367.4+5.1
- DocVQA-valMinistral 3 3B by 1.987.789.6+1.9
- MMMBMinistral 3 3B by 1.771.773.4+1.7
- MMBench DEV_EN_V11tie69.669.2tie
- MMBench_DEVtie0.60.7tie
- OCRBench v2_entie41.541.4tie
- AA Intelligencetie4.84.8tie
Questions people ask
Which is better, LFM2.5-VL-1.6B or Ministral 3 3B?
Ministral 3 3B wins four of the seven areas where both have results: coding, reasoning, facts and long documents. LFM2.5-VL-1.6B wins following instructions. They are level on agents and images and charts.
Which is better for coding?
Ministral 3 3B. It wins 1 of the 2 coding tests both models report; LFM2.5-VL-1.6B wins none, and 1 is a tie.
How do you compare the two?
We use the 27 benchmark tests both models have published scores on. The verdict counts the 12 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 15 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.