Scores updated · 6 tests both models report · How we compare
Every test, side by side
All 6 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results6 tests
- GSM8KArctic Instruct by 31.676.845.2+31.6
- MATHArctic Instruct by 28.351.923.6+28.3
- NarrativeQACommand by 9.565.374.9+9.5
- OpenBookQAArctic Instruct by 5.482.877.4+5.4
- MMLUArctic Instruct by 557.552.5+5
- naturalquestions_closedbooktie3939.1tie
Questions people ask
How do you compare the two?
We use the 6 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.