Across 31 shared benchmarks, Gemma 3 1B Instruct scores higher on 3 and LFM2.5-1.2B-Instruct on 25, with 3 level. The widest gap is humaneval, where Gemma 3 1B Instruct scores 41.5 against 5.3.
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.