Claude 3 Opus vs o1
Wins 0 of 4 areas
—
Wins 4 of 4 areas
Coding · Reasoning · Facts · Images and charts
o1 is the stronger all-rounder.
Scores updated · 22 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Coding, facts and images and charts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
o1 costs 17% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where o1 pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+28.5points ahead
- Graduate-level biology, physics and chemistry questionsGPQA Diamond+25.8points ahead
- Math problems shown in pictures and chartsMathVista+21.3points ahead
Where Claude 3 Opus pulls ahead
No clear win on a test scored out of 100.
Every test, side by side
All 22 tests both models report. The winning score is in its model's colour; marks a score checked independently.
Reasoningo1
- GPQA Diamondo1 by 25.848.974.7+25.8
- SimpleBencho1 by 16.623.540.1+16.6
- Humanity's Last Examo1 by 4.22.87+4.2
Other results16 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- MATHo1 by 36.360.196.4+36.3
- metr_hcasto1 by 2220.642.6+22
- Artificial Analysis Coding Indexo1 by 20.219.539.7+20.2
- AA Intelligenceo1 by 6.58.715.2+6.5
- MMLUo1 by 586.891.8+5
- XSTesto1 by 4.592.597+4.5
- AIR-Bench 2024Claude 3 Opus by 4.484.480+4.4
- bbqo1 by 3.39497.3+3.3
- HumanEvalo1 by 3.284.988.1+3.2
- GSM8Ko1 by 2.19597.1+2.1
- anthropic_red_teamClaude 3 Opus by 1.599.898.3+1.5
- metr_swaao1 by 1.598.5100+1.5
- HarmBenchClaude 3 Opus by 1.197.496.3+1.1
- simple_safety_testsClaude 3 Opus by 110099+1
- MGSMtie90.790.8tie
- metr_re_benchtie00tie
Questions people ask
Which is better, Claude 3 Opus or o1?
o1 wins all four areas where both have results: coding, reasoning, facts and images and charts. Claude 3 Opus wins none.
Which is better for coding?
o1. It wins the one coding test both models report.
Which is cheaper?
Claude 3 Opus costs $15.00 per million input tokens and $75.00 per million output tokens; o1 costs $15.00 and $60.00. That makes o1 about 17% cheaper for the same work.
How do you compare the two?
We use the 22 benchmark tests both models have published scores on. The verdict counts the 6 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 16 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.