Scores updated · 6 tests both models report · How we compare
Every test, side by side
All 6 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results6 tests
- OpenBookQACommand by 48.877.428.6+48.8
- GSM8KCommand by 42.445.22.8+42.4
- MMLUCommand by 28.252.524.3+28.2
- MATHCommand by 2123.62.6+21
- naturalquestions_closedbookCommand by 19.439.119.7+19.4
- NarrativeQACommand by 11.674.963.3+11.6
Questions people ask
How do you compare the two?
We use the 6 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.