Scores updated · 6 tests both models report · How we compare
Every test, side by side
All 6 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Other results6 tests
- HarmBenchClaude 3 Sonnet by 40.295.855.6+40.2
- AIR-Bench 2024Claude 3 Sonnet by 26.984.757.8+26.9
- XSTestPalmyra Med by 10.585.896.3+10.5
- bbqClaude 3 Sonnet by 9.99080.1+9.9
- anthropic_red_teamClaude 3 Sonnet by 299.897.8+2
- simple_safety_testsClaude 3 Sonnet by 1.510098.5+1.5
Questions people ask
How do you compare the two?
We use the 6 benchmark tests both models have published scores on. None of them is in the eight capability areas we count, so this page lists them without a verdict. Each score is the one shown on the model's own page.