Scores updated · 9 tests both models report · How we compare
Every test, side by side
All 9 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results9 tests
- JHumanEval+LFM2.5-1.2B-JP-202606 by 20.728.749.4+20.7
- Domain AvgLFM2.5-1.2B-JP-202606 by 13.439.753.1+13.4
- J-MATH500LFM2.5-1.2B-JP-202606 by 12.85062.8+12.8
- Avg.LFM2.5-1.2B-JP-202606 by 12.450.162.5+12.4
- J-GSM8KLFM2.5-1.2B-JP-202606 by 1250.262.2+12
- JMMLULFM2.5-1.2B-JP-202606 by 6.647.654.2+6.6
- JMMLU-ProXLFM2.5-1.2B-JP-202606 by 4.831.436.2+4.8
- JGPQALFM2.5-1.2B-Instruct by 331.728.7+3
- J-BFCLv3LFM2.5-1.2B-JP-202606 by 1.746.348+1.7
Questions people ask
How do you compare the two?
We use the 9 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.