Scores updated · 9 tests both models report · How we compare
Every test, side by side
All 9 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results9 tests
- J-GSM8KLFM2.5-1.2B-JP-202606 by 16.262.246+16.2
- Avg.LFM2.5-1.2B-JP-202606 by 11.362.551.2+11.3
- Domain AvgLFM2.5-1.2B-JP-202606 by 8.353.144.8+8.3
- JMMLULFM2.5-1.2B-JP-202606 by 6.554.247.7+6.5
- J-MATH500LFM2.5-1.2B-JP-202606 by 6.462.856.4+6.4
- JMMLU-ProXLFM2.5-1.2B-JP-202606 by 5.436.230.8+5.4
- J-BFCLv3Qwen3 1.7B (instruct) by 4.54852.5+4.5
- JGPQALFM2.5-1.2B-JP-202606 by 2.428.726.3+2.4
- JHumanEval+LFM2.5-1.2B-JP-202606 by 1.849.447.6+1.8
Questions people ask
How do you compare the two?
We use the 9 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.