Gemma 3 4B vs LFM2-2.6B
Wins 3 of 6 areas
Coding · Long documents · Following instructions
Wins 1 of 6 areas
Reasoning
Gemma 3 4B wins more areas, narrowly.LFM2-2.6B is better at reasoning.
Scores updated · 25 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software20Gemma 3 4B2 of 3 tests · 1 tie
- Long documentsFinding answers in very long texts10Gemma 3 4B1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemma 3 4B1 of 1 test
- ReasoningHard problems that need careful thinking01LFM2-2.6B1 of 3 tests · 2 ties
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
Agents, long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Gemma 3 4B pulls ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+8.8points ahead
- Code for real scientific research problemsSciCode+4.8points ahead
- Recent programming contest problemsLiveCodeBench v6+4.5points ahead
Where LFM2-2.6B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+35.3points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+1.5points ahead
Every test, side by side
All 25 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemma 3 4B
- SciCodeGemma 3 4B by 4.87.32.5+4.8
- LiveCodeBench v6Gemma 3 4B by 4.518.914.4+4.5
- Terminal-Bench Hardtie0.80.8tie
ReasoningLFM2-2.6B
- GPQA DiamondLFM2-2.6B by 1.529.130.6+1.5
- Humanity's Last Examtie5.35.5tie
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationLFM2-2.6B by 35.31.837.1+35.3
- AA-Omniscience · AccuracyGemma 3 4B by 2.37.75.4+2.3
Following instructionsGemma 3 4B
- IFBenchGemma 3 4B by 8.828.319.5+8.8
Other results14 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceLFM2-2.6B by 28.8-82.9-54.1+28.8
- MMLU-ProGemma 3 4B by 17.643.626+17.6
- MGSMGemma 3 4B by 1387.374.3+13
- IFEvalGemma 3 4B by 10.690.279.6+10.6
- MATH-500 (EM)Gemma 3 4B by 9.673.263.6+9.6
- τ²-Bench Telecom (AA run)LFM2-2.6B by 8.5513.5+8.5
- GSMPlusGemma 3 4B by 7.668.460.8+7.6
- GSM8KGemma 3 4B by 6.889.282.4+6.8
- MMLULFM2-2.6B by 658.464.4+6
- MMMLULFM2-2.6B by 5.350.155.4+5.3
- HumanEval+Gemma 3 4B by 4.962.857.9+4.9
- LCB v5Gemma 3 4B by 4.719.114.4+4.7
- AA Agentic IndexLFM2-2.6B by 2.81.74.5+2.8
- AA Intelligencetie4.85.2tie
Questions people ask
Which is better, Gemma 3 4B or LFM2-2.6B?
Gemma 3 4B wins three of the six areas where both have results: coding, long documents and following instructions. LFM2-2.6B wins reasoning. They are level on agents and facts.
Which is better for coding?
Gemma 3 4B. It wins 2 of the 3 coding tests both models report; LFM2-2.6B wins none, and 1 is a tie.
How do you compare the two?
We use the 25 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 14 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.