Scores updated · 9 tests both models report · How we compare
Every test, side by side
All 9 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results9 tests
- J-MATH500LFM2.5-1.2B-JP-202606 by 47.215.662.8+47.2
- Avg.LFM2.5-1.2B-JP-202606 by 37.924.662.5+37.9
- J-BFCLv3LFM2.5-1.2B-JP-202606 by 30.717.348+30.7
- Domain AvgLFM2.5-1.2B-JP-202606 by 29.223.953.1+29.2
- J-GSM8KLFM2.5-1.2B-JP-202606 by 28.633.662.2+28.6
- JHumanEval+LFM2.5-1.2B-JP-202606 by 24.42549.4+24.4
- JMMLU-ProXLFM2.5-1.2B-JP-202606 by 22.114.136.2+22.1
- JMMLULFM2.5-1.2B-JP-202606 by 19.734.554.2+19.7
- JGPQALFM2.5-1.2B-JP-202606 by 4.524.228.7+4.5
Questions people ask
How do you compare the two?
We use the 9 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.