Scores updated · 7 tests both models report · How we compare
Every test, side by side
All 7 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results7 tests
- NarrativeQALlama (65B) by 64.411.175.5+64.4
- GSM8KClaude 3 Sonnet by 43.492.348.9+43.4
- naturalquestions_closedbookLlama (65B) by 40.52.843.3+40.5
- MMLUClaude 3 Sonnet by 19.978.358.4+19.9
- MATHClaude 3 Sonnet by 17.443.125.7+17.4
- OpenBookQAClaude 3 Sonnet by 16.491.875.4+16.4
- AA Intelligencetie5.95tie
Questions people ask
How do you compare the two?
We use the 7 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.