LFM2-2.6B vs Qwen3 1.7B
Wins 0 of 6 areas
—
Wins 3 of 6 areas
Coding · Reasoning · Following instructions
Qwen3 1.7B wins more areas, narrowly.
Scores updated · 24 tests both models report · How we compare
Where each one wins
Tests won in each of the six areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software02Qwen3 1.7B2 of 3 tests · 1 tie
- ReasoningHard problems that need careful thinking01Qwen3 1.7B1 of 3 tests · 2 ties
- Following instructionsDoing exactly what it is asked01Qwen3 1.7B1 of 1 test
- AgentsCarrying out multi-step tasks on its own00Even0 each · 1 tie
- FactsGetting facts right instead of making them up11Even1 each
- Long documentsFinding answers in very long texts00Even0 each · 1 tie
Agents, long documents and following instructions rest on a single test each.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Qwen3 1.7B pulls ahead
- Recent programming contest problemsLiveCodeBench v6+9.7points ahead
- Follows unfamiliar, precisely checkable instructionsIFBench+7.4points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+5points ahead
Where LFM2-2.6B pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+31.9points ahead
Every test, side by side
All 24 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingQwen3 1.7B
- LiveCodeBench v6Qwen3 1.7B by 9.714.424.1+9.7
- SciCodeQwen3 1.7B by 1.82.54.3+1.8
- Terminal-Bench Hardtie0.80tie
ReasoningQwen3 1.7B
- GPQA DiamondQwen3 1.7B by 530.635.6+5
- Humanity's Last Examtie5.54.6tie
- CritPttie00tie
FactsEven
- AA-Omniscience · Non-hallucinationLFM2-2.6B by 31.937.15.2+31.9
- AA-Omniscience · AccuracyQwen3 1.7B by 3.55.48.9+3.5
Following instructionsQwen3 1.7B
- IFBenchQwen3 1.7B by 7.419.526.9+7.4
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GSM8KLFM2-2.6B by 3182.451.4+31
- MMLU-ProQwen3 1.7B by 312657+31
- AA-OmniscienceLFM2-2.6B by 23.4-54.1-77.5+23.4
- MATH-500 (EM)Qwen3 1.7B by 18.363.681.9+18.3
- τ²-Bench Telecom (AA run)Qwen3 1.7B by 12.513.526+12.5
- LCB v5Qwen3 1.7B by 12.114.426.5+12.1
- MMMLULFM2-2.6B by 8.955.446.5+8.9
- MGSMLFM2-2.6B by 7.774.366.6+7.7
- IFEvalLFM2-2.6B by 5.679.674+5.6
- MMLULFM2-2.6B by 5.364.459.1+5.3
- AA Agentic IndexQwen3 1.7B by 4.24.58.7+4.2
- HumanEval+Qwen3 1.7B by 3.157.961+3.1
- AA Intelligencetie5.25.2tie
Questions people ask
Which is better, LFM2-2.6B or Qwen3 1.7B?
Qwen3 1.7B wins three of the six areas where both have results: coding, reasoning and following instructions. LFM2-2.6B wins none. They are level on agents, facts and long documents.
Which is better for coding?
Qwen3 1.7B. It wins 2 of the 3 coding tests both models report; LFM2-2.6B wins none, and 1 is a tie.
How do you compare the two?
We use the 24 benchmark tests both models have published scores on. The verdict counts the 11 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.