Scores updated · 9 tests both models report · How we compare
Every test, side by side
All 9 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results9 tests
- J-MATH500LFM2.5-1.2B-JP-202606 by 37.425.462.8+37.4
- Avg.LFM2.5-1.2B-JP-202606 by 28.434.162.5+28.4
- JMMLU-ProXLFM2.5-1.2B-JP-202606 by 20.915.336.2+20.9
- JMMLULFM2.5-1.2B-JP-202606 by 20.333.954.2+20.3
- J-GSM8KLFM2.5-1.2B-JP-202606 by 19.442.862.2+19.4
- Domain AvgLFM2.5-1.2B-JP-202606 by 14.638.553.1+14.6
- JGPQALFM2.5-1.2B-JP-202606 by 4.324.428.7+4.3
- J-BFCLv3Granite 4.0 1B by 2.650.648+2.6
- JHumanEval+Granite 4.0 1B by 1.851.249.4+1.8
Questions people ask
How do you compare the two?
We use the 9 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.