Claude Opus 4.7 vs Claude Opus 4.8
Wins 2 of 8 areas
Images and charts · Long documents
Wins 6 of 8 areas
Coding · Agents · Reasoning · Facts · Math · Following instructions
Claude Opus 4.8 is the stronger all-rounder.Claude Opus 4.7 is better at images and charts.
Both rank among the ten best models we track in coding.
Scores updated · 69 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- MathCompetition and research-level math06Claude Opus 4.86 of 6 tests
- CodingWriting and fixing software05Claude Opus 4.85 of 9 tests · 4 ties
- AgentsCarrying out multi-step tasks on its own05Claude Opus 4.85 of 6 tests · 1 tie
- ReasoningHard problems that need careful thinking15Claude Opus 4.85 of 7 tests · 1 tie
- FactsGetting facts right instead of making them up02Claude Opus 4.82 of 3 tests · 1 tie
- Following instructionsDoing exactly what it is asked02Claude Opus 4.82 of 2 tests
- Images and chartsUnderstanding pictures, charts and video20Claude Opus 4.72 of 2 tests
- Long documentsFinding answers in very long texts10Claude Opus 4.71 of 1 test
Long documents rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Claude Opus 4.8 pulls ahead
- Proof problems from the US Math OlympiadUSAMO 2026+27.4points ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+24.4points ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+9.8points ahead
Where Claude Opus 4.7 pulls ahead
- Abstract visual puzzles that people can solveARC-AGI-2+3.7points ahead
- Reasoning about charts from research papersCharXiv (RQ)+1.1points ahead
- Reasons across sets of long documentsAA-LCR+1points ahead
Every test, side by side
All 69 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Opus 4.8
- Terminal-Bench HardClaude Opus 4.8 by 6.851.558.3+6.8
- SWE-bench ProClaude Opus 4.8 by 4.964.369.2+4.9
- SWE-bench MultilingualClaude Opus 4.8 by 3.980.584.4+3.9
- Terminal-Bench 2.1Claude Opus 4.8 by 1.583.184.6+1.5
- SWE-bench VerifiedClaude Opus 4.8 by 187.688.6+1
- LiveBench · Codingtie82.181.8tie
- LiveBench · Agentic Codingtie50.750.5tie
- LMArena · WebDevtie15581556tie
- SciCodetie54.554.4tie
AgentsClaude Opus 4.8
- AA ApexAgentsClaude Opus 4.8 by 5.533.939.4+5.5
- OSWorld-VerifiedClaude Opus 4.8 by 5.47883.4+5.4
- BrowseCompClaude Opus 4.8 by 579.384.3+5
- GDPValClaude Opus 4.8 by 542.847.8+5
- MCP AtlasClaude Opus 4.8 by 4.977.382.2+4.9
- τ-Bench V3 · Bankingtie34.634.2tie
ReasoningClaude Opus 4.8
- CritPtClaude Opus 4.8 by 8.91220.9+8.9
- Humanity's Last ExamClaude Opus 4.8 by 6.442.348.7+6.4
- ARC-AGI-2Claude Opus 4.7 by 3.775.872.1+3.7
- SimpleBenchClaude Opus 4.8 by 3.161.764.8+3.1
- LiveBench · ReasoningClaude Opus 4.8 by 287.289.2+2
- ARC-AGI-3Claude Opus 4.8 by 1.30.21.5+1.3
- GPQA Diamondtie91.492tie
FactsClaude Opus 4.8
- AA-Omniscience · Non-hallucinationClaude Opus 4.8 by 357.760.7+3
- SimpleQA VerifiedClaude Opus 4.8 by 1.351.753+1.3
- AA-Omniscience · Accuracytie48.948.8tie
Images and chartsClaude Opus 4.7
- LMArena · VisionClaude Opus 4.7 by 18 rating points13131295+18 rating
- CharXiv (RQ)Claude Opus 4.7 by 1.19189.9+1.1
MathClaude Opus 4.8
- USAMO 2026Claude Opus 4.8 by 27.469.396.7+27.4
- FrontierMath Tier 4Claude Opus 4.8 by 24.431.756.1+24.4
- FrontierMath Tiers 1-3 (v2)Claude Opus 4.8 by 9.870.280+9.8
- AIME 2026Claude Opus 4.8 by 4.295.8100+4.2
- HMMT Feb. 2026Claude Opus 4.8 by 1.693.995.5+1.6
- LiveBench · MathematicsClaude Opus 4.8 by 1.492.894.3+1.4
Following instructionsClaude Opus 4.8
- LiveBench · Instruction FollowingClaude Opus 4.8 by 5.366.772+5.3
- IFBenchClaude Opus 4.8 by 3.658.662.2+3.6
Other results33 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- OfficeQA ProClaude Opus 4.7 by 14.480.666.2+14.4
- GDPval-AA (Elo)Claude Opus 4.8 by 137 rating points17531890+137 rating
- livebench_data_analysisClaude Opus 4.7 by 12.378.366+12.3
- Finance AgentClaude Opus 4.7 by 10.564.453.9+10.5
- Blueprint-Bench 2Claude Opus 4.7 by 1024.514.5+10
- GraphWalks BFS 256KClaude Opus 4.8 by 976.985.9+9
- OfficeQAClaude Opus 4.7 by 8.786.377.6+8.7
- frontiermath_tier_4_v1Claude Opus 4.8 by 8.422.931.3+8.4
- ScreenSpot-Pro (No tools)Claude Opus 4.8 by 8.479.587.9+8.4
- HiL-Bench (Tools-allowed)Claude Opus 4.7 by 6.441.735.3+6.4
- τ²-Bench Telecom (AA run)Claude Opus 4.8 by 5.888.694.4+5.8
- CyberGymClaude Opus 4.8 by 5.773.178.8+5.7
- GraphWalks — Parents task, 256K-token context subsetClaude Opus 4.8 by 5.793.699.3+5.7
- Terminal-Bench 2.0Claude Opus 4.8 by 5.269.474.6+5.2
- Toolathlon Pass@∞Claude Opus 4.7 by 4.752.848.1+4.7
- DeepSWEClaude Opus 4.8 by 45458+4
- SWE-bench MultimodalClaude Opus 4.8 by 3.934.538.4+3.9
- HLE (with tools)Claude Opus 4.8 by 3.254.757.9+3.2
- AA Agentic IndexClaude Opus 4.8 by 3.139.542.6+3.1
- Legal Agent BenchmarkClaude Opus 4.8 by 2.97.110+2.9
- ChartQAPro (With tools)Claude Opus 4.8 by 2.569.872.3+2.5
- Finance Agent v2Claude Opus 4.8 by 2.451.553.9+2.4
- ChartQAPro (No tools)Claude Opus 4.8 by 1.867.669.4+1.8
- livebench_languageClaude Opus 4.8 by 1.877.979.7+1.8
- AA-OmniscienceClaude Opus 4.8 by 1.527.328.8+1.5
- Toolathlon Avg turnsClaude Opus 4.7 by 1.425.924.5+1.4
- AA IntelligenceClaude Opus 4.8 by 1.140.741.8+1.1
- CharXiv Reasoning (With tools)Claude Opus 4.7 by 1.19189.9+1.1
- Artificial Analysis Coding Indextie73.674.3tie
- Toolathlontie59.359.9tie
- ARC-AGI-1tie9292.5tie
- LiveBenchtie76.977.2tie
- ScreenSpot-Pro (With tools)tie87.687.9tie
Questions people ask
Which is better, Claude Opus 4.7 or Claude Opus 4.8?
Claude Opus 4.8 wins six of the eight areas we test: coding, agents, reasoning, facts, math and following instructions. Claude Opus 4.7 wins images and charts and long documents.
Which is better for coding?
Claude Opus 4.8. It wins 5 of the 9 coding tests both models report; Claude Opus 4.7 wins none, and 4 are ties.
Which is cheaper?
Claude Opus 4.7 costs $5.00 per million input tokens and $25.00 per million output tokens; Claude Opus 4.8 costs $5.00 and $25.00. For a million tokens read plus a million written, they cost about the same.
How do you compare the two?
We use the 69 benchmark tests both models have published scores on. The verdict counts the 36 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 33 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.