Scores updated · 6 tests both models report · How we compare
Every test, side by side
All 6 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results6 tests
- AIR-Bench 2024Claude 3 Sonnet by 18.484.766.3+18.4
- HarmBenchClaude 3 Sonnet by 11.795.884.1+11.7
- XSTestPalmyra Fin by 10.485.896.2+10.4
- bbqPalmyra Fin by 4.29094.2+4.2
- anthropic_red_teamtie99.899.5tie
- simple_safety_teststie100100tie
Questions people ask
How do you compare the two?
We use the 6 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.