o1 is the stronger all-rounder.Claude 3.5 Sonnet is cheaper.
Scores updated · 35 tests both models report · How we compare
Where each one wins
Tests won in each of the three areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Claude 3.5 Sonnet costs 76% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where o1 pulls ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+18.7points ahead
- Common-sense trick questionsSimpleBench+12.6points ahead
- College exam questions with charts, maps and diagramsMMMU+9.3points ahead
Where Claude 3.5 Sonnet pulls ahead
- Fixes real GitHub issues in Python projectsSWE-bench Verified+7.7points ahead
Every test, side by side
All 35 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingEven
- SWE-bench VerifiedClaude 3.5 Sonnet by 7.74941.3+7.7
- SciCodeo1 by 4.231.635.8+4.2
Reasoningo1
- GPQA Diamondo1 by 18.75674.7+18.7
- SimpleBencho1 by 12.627.540.1+12.6
- Humanity's Last Examo1 by 3.83.27+3.8
Images and chartso1
Other results28 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- Codeforceso1 by 1344 rating points7172061+1344 rating
- Codeforces (Rating)o1 by 1344 rating points7172061+1344 rating
- AIME 2025o1 by 78.43.381.7+78.4
- Codeforces (Percentile)o1 by 76.320.396.6+76.3
- AIME 2024o1 by 63.21679.2+63.2
- LiveCodeBench (Pass@1-COT)o1 by 29.633.863.4+29.6
- MATHo1 by 25.371.196.4+25.3
- MATH-500 (EM)o1 by 18.178.396.4+18.1
- Aider-Polygloto1 by 16.445.361.7+16.4
- SimpleQAo1 by 1428.442.4+14
- Artificial Analysis Coding Indexo1 by 13.72639.7+13.7
- AA Intelligenceo1 by 87.215.2+8
- AIR-Bench 2024Claude 3.5 Sonnet by 5.985.980+5.9
- HumanEvalClaude 3.5 Sonnet by 5.693.788.1+5.6
- metr_re_benchClaude 3.5 Sonnet by 550+5
- TAU-bench (airline)o1 by 44650+4
- MMLUo1 by 3.188.791.8+3.1
- bbqo1 by 2.494.997.3+2.4
- DROP (3-shot F1)o1 by 1.988.390.2+1.9
- metr_hcasto1 by 1.940.742.6+1.9
- HarmBenchClaude 3.5 Sonnet by 1.898.196.3+1.8
- TAU-bench (retail)o1 by 1.669.270.8+1.6
- anthropic_red_teamClaude 3.5 Sonnet by 1.599.898.3+1.5
- XSTesto1 by 1.495.697+1.4
- metr_swaao1 by 1.398.7100+1.3
- simple_safety_testsClaude 3.5 Sonnet by 110099+1
- MGSMtie91.690.8tie
- GSM8Ktie96.497.1tie
Questions people ask
Which is better, Claude 3.5 Sonnet or o1?
o1 wins two of the three areas where both have results: reasoning and images and charts. Claude 3.5 Sonnet wins none, but costs 76% less. They are level on coding.
Which is better for coding?
Neither. They win 1 coding test each of the 2 both models report.
Which is cheaper?
Claude 3.5 Sonnet costs $3.00 per million input tokens and $15.00 per million output tokens; o1 costs $15.00 and $60.00. That makes Claude 3.5 Sonnet about 76% cheaper for the same work.
How do you compare the two?
We use the 35 benchmark tests both models have published scores on. The verdict counts the 7 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 28 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.