Scores updated · 6 tests both models report · How we compare
Every test, side by side
All 6 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results6 tests
- naturalquestions_closedbookCommand by 23.639.115.5+23.6
- GSM8KPhi 2 by 12.945.258.1+12.9
- NarrativeQACommand by 4.674.970.3+4.6
- OpenBookQAPhi 2 by 2.477.479.8+2.4
- MATHPhi 2 by 1.923.625.5+1.9
- MMLUtie52.551.8tie
Questions people ask
How do you compare the two?
We use the 6 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.