Claude Mythos 5 wins more areas, narrowly.Claude Opus 5 is cheaper and better at coding.
Scores updated · 26 tests both models report · How we compare
Where each one wins
Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
Agents and images and charts rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Claude Opus 5 costs 50% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Claude Mythos 5 pulls ahead
- Reasoning about charts from research papersCharXiv (RQ)+5.2points ahead
- Very hard expert questions across many subjectsHumanity's Last Exam+2.9points ahead
- Completes tasks by operating a computer desktopOSWorld-Verified+1.6points ahead
Where Claude Opus 5 pulls ahead
- Fixes real GitHub issues in many programming languagesSWE-bench Multilingual+2.9points ahead
- Abstract visual puzzles that people can solveARC-AGI-2+1.2points ahead
- Command-line tasks in a real terminalTerminal-Bench 2.1+1.1points ahead
Every test, side by side
All 26 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Opus 5
- SWE-bench MultilingualClaude Opus 5 by 2.986.689.5+2.9
- Terminal-Bench 2.1Claude Opus 5 by 1.18889.1+1.1
- SWE-bench Protie8079.2tie
- SWE-bench Verifiedtie95.596tie
ReasoningEven
- Humanity's Last ExamClaude Mythos 5 by 2.957.854.9+2.9
- ARC-AGI-2Claude Opus 5 by 1.289.290.4+1.2
- CritPttie28.629.1tie
Images and chartsClaude Mythos 5
- CharXiv (RQ)Claude Mythos 5 by 5.288.983.7+5.2
Other results17 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDPval-AA v2Claude Opus 5 by 101 rating points17231824+101 rating
- Terminal-Bench 4.0Claude Opus 5 by 9.84251.8+9.8
- SWE-bench MultimodalClaude Opus 5 by 5.354.159.4+5.3
- Terminal-Bench-Science 0.1Claude Opus 5 by 5.324.730+5.3
- HealthBench (raw)Claude Opus 5 by 4.662.567.1+4.6
- HealthBench ProfessionalClaude Mythos 5 by 3.563.359.8+3.5
- OSWorld 2.0 (strict)Claude Opus 5 by 3.536.139.6+3.5
- HealthBench Professional (raw)Claude Opus 5 by 3.170.373.4+3.1
- OSWorld 2.0 (partial)Claude Opus 5 by 2.572.975.4+2.5
- ProteinGym HardClaude Opus 5 by 1.945.847.7+1.9
- GDP (Surge AI)Claude Mythos 5 by 1.887.385.5+1.8
- Toolathlon VerifiedClaude Opus 5 by 1.379.380.6+1.3
- ARC-AGI-1Claude Mythos 5 by 198.597.5+1
- OfficeQAtie7978.1tie
- Program Benchtie86.385.4tie
- HLE (with tools)tie63.863.6tie
- OfficeQA Protie67.166.9tie
Questions people ask
Which is better, Claude Mythos 5 or Claude Opus 5?
Claude Mythos 5 wins two of the four areas where both have results: agents and images and charts. Claude Opus 5 wins coding, and costs 50% less. They are level on reasoning.
Which is better for coding?
Claude Opus 5. It wins 2 of the 4 coding tests both models report; Claude Mythos 5 wins none, and 2 are ties.
Which is cheaper?
Claude Mythos 5 costs $10.00 per million input tokens and $50.00 per million output tokens; Claude Opus 5 costs $5.00 and $25.00. That makes Claude Opus 5 about 50% cheaper for the same work.
How do you compare the two?
We use the 26 benchmark tests both models have published scores on. The verdict counts the 9 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 17 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.