LFM2.5-2.6B vs LFM2.5-8B-A1B
Wins 4 of 6 areas
Coding · Reasoning · Long documents · Following instructions
Wins 0 of 6 areas
—
LFM2.5-2.6B is the stronger all-rounder.
Scores updated · 20 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking10LFM2.5-2.6B1 of 3 tests · 2 ties
- CodingWriting and fixing software10LFM2.5-2.6B1 of 1 test
- Long documentsFinding answers in very long texts10LFM2.5-2.6B1 of 1 test
- Following instructionsDoing exactly what it is asked10LFM2.5-2.6B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
Coding, agents, long documents and following instructions rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where LFM2.5-2.6B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+30.9points ahead
- Code for real scientific research problemsSciCode+6.6points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+4.5points ahead
Where LFM2.5-8B-A1B pulls ahead
- Answers hard knowledge questions correctlyAA-Omniscience · Accuracy+4.9points ahead
Every test, side by side
All 20 tests both models report. The winning score is in its model's colour; marks a score checked independently.
ReasoningLFM2.5-2.6B
- GPQA DiamondLFM2.5-2.6B by 4.555.851.3+4.5
- Humanity's Last Examtie6.26.9tie
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationLFM2.5-2.6B by 30.98453.1+30.9
- AA-Omniscience · AccuracyLFM2.5-8B-A1B by 4.94.39.3+4.9
Following instructionsLFM2.5-2.6B
- IFBenchLFM2.5-2.6B by 3.659.255.6+3.6
Other results11 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceLFM2.5-2.6B by 22.3-10.9-33.2+22.3
- AIME25 no toolsLFM2.5-2.6B by 9.451.942.5+9.4
- BFCLv4LFM2.5-2.6B by 7.256.949.7+7.2
- MT-BenchLFM2.5-8B-A1B by 3.45.18.5+3.4
- AA Agentic IndexLFM2.5-8B-A1B by 32.45.4+3
- MATH500 (Pass@1)LFM2.5-8B-A1B by 2.95.48.3+2.9
- HumanEvalLFM2.5-8B-A1B by 2.54.57+2.5
- MBPPLFM2.5-8B-A1B by 2.24.76.9+2.2
- Artificial Analysis Coding IndexLFM2.5-2.6B by 2.17.75.6+2.1
- AA IntelligenceLFM2.5-2.6B by 1.28.47.2+1.2
- GSM8Ktie4.34tie
Questions people ask
Which is better, LFM2.5-2.6B or LFM2.5-8B-A1B?
LFM2.5-2.6B wins four of the six areas where both have results: coding, reasoning, long documents and following instructions. LFM2.5-8B-A1B wins none. They are level on agents and facts.
Which is better for coding?
LFM2.5-2.6B. It wins the one coding test both models report.
How do you compare the two?
We use the 20 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 11 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.