Claude Mythos 5 vs Claude Opus 4.8
Wins 4 of 5 areas
Coding · Agents · Reasoning · Math
Wins 1 of 5 areas
Images and charts
Claude Mythos 5 is the stronger all-rounder.Claude Opus 4.8 is cheaper and better at images and charts.
Scores updated · 34 tests both models report · How we compare
Where each one wins
Tests won in each of the five areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software40Claude Mythos 54 of 4 tests
- ReasoningHard problems that need careful thinking30Claude Mythos 53 of 3 tests
- AgentsCarrying out multi-step tasks on its own10Claude Mythos 51 of 1 test
- MathCompetition and research-level math10Claude Mythos 51 of 1 test
- Images and chartsUnderstanding pictures, charts and video01Claude Opus 4.81 of 1 test
Agents, images and charts and math rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Claude Opus 4.8 costs 50% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude Mythos 5 pulls ahead
- Abstract visual puzzles that people can solveARC-AGI-2+17.1points ahead
- Long, multi-file coding tasks in real codebasesSWE-bench Pro+10.8points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+9.1points ahead
Where Claude Opus 4.8 pulls ahead
- Reasoning about charts from research papersCharXiv (RQ)+1points ahead
Every test, side by side
All 34 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Mythos 5
- SWE-bench ProClaude Mythos 5 by 10.88069.2+10.8
- SWE-bench VerifiedClaude Mythos 5 by 6.995.588.6+6.9
- Terminal-Bench 2.1Claude Mythos 5 by 3.48884.6+3.4
- SWE-bench MultilingualClaude Mythos 5 by 2.286.684.4+2.2
ReasoningClaude Mythos 5
- ARC-AGI-2Claude Mythos 5 by 17.189.272.1+17.1
- Humanity's Last ExamClaude Mythos 5 by 9.157.848.7+9.1
- CritPtClaude Mythos 5 by 7.728.620.9+7.7
Images and chartsClaude Opus 4.8
- CharXiv (RQ)Claude Opus 4.8 by 188.989.9+1
Other results24 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- RiemannBenchClaude Mythos 5 by 215534+21
- Terminal-Bench 4.0Claude Mythos 5 by 18.44223.6+18.4
- SWE-bench MultimodalClaude Mythos 5 by 15.754.138.4+15.7
- Program BenchClaude Mythos 5 by 14.486.371.9+14.4
- GDPval-AA v2Claude Mythos 5 by 130 rating points17231593+130 rating
- HealthBench Professional (raw)Claude Mythos 5 by 1070.360.3+10
- HealthBench ProfessionalClaude Mythos 5 by 7.563.355.8+7.5
- Toolathlon Pass@∞Claude Mythos 5 by 7.555.648.1+7.5
- Legal Agent Benchmark, Full Public SetClaude Mythos 5 by 7.316.99.6+7.3
- Toolathlon Avg turnsClaude Opus 4.8 by 6.917.624.5+6.9
- ARC-AGI-1Claude Mythos 5 by 698.592.5+6
- HLE (with tools)Claude Mythos 5 by 5.963.857.9+5.9
- GraphWalks BFS 256KClaude Mythos 5 by 5.291.185.9+5.2
- BrowseComp (multi-agent)Claude Mythos 5 by 4.893.388.5+4.8
- BrowseComp (single-agent)Claude Mythos 5 by 3.78884.3+3.7
- HealthBench (raw)Claude Mythos 5 by 3.762.558.8+3.7
- CharXiv Reasoning (With tools)Claude Mythos 5 by 3.693.589.9+3.6
- HealthBenchClaude Mythos 5 by 3.462.759.3+3.4
- GDP (Surge AI)Claude Mythos 5 by 2.587.384.8+2.5
- OfficeQAClaude Mythos 5 by 1.47977.6+1.4
- ToolathlonClaude Mythos 5 by 1.261.159.9+1.2
- OfficeQA Protie67.166.2tie
- GraphWalks — Parents task, 256K-token context subsettie10099.3tie
- Toolathlon Verifiedtie79.379.9tie
Questions people ask
Which is better, Claude Mythos 5 or Claude Opus 4.8?
Claude Mythos 5 wins four of the five areas where both have results: coding, agents, reasoning and math. Claude Opus 4.8 wins images and charts, and costs 50% less.
Which is better for coding?
Claude Mythos 5. It wins 4 of the 4 coding tests both models report; Claude Opus 4.8 wins none.
Which is cheaper?
Claude Mythos 5 costs $10.00 per million input tokens and $50.00 per million output tokens; Claude Opus 4.8 costs $5.00 and $25.00. That makes Claude Opus 4.8 about 50% cheaper for the same work.
How do you compare the two?
We use the 34 benchmark tests both models have published scores on. The verdict counts the 10 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 24 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.