Claude Mythos Preview vs Claude Opus 4.6
15 SHARED BENCHMARKSAcross 15 shared benchmarks, Claude Mythos Preview scores higher on 15 and Claude Opus 4.6 on 0. The widest gap is SWE-bench Multimodal, where Claude Mythos Preview scores 59 against 27.1.
ANTHROPICVSANTHROPIC15 SHARED15–0 HEAD-TO-HEADNEWEST SCORE ADDED
At a glance
MakerAnthropicAnthropic
Released—
Price per 1M tokens input / output—$5.00 / $25.00
Cost of 1M in + 1M out—$30.00
Head-to-head of 15 shared benchmarks15 wins0 wins
Scores tracked independently verified25 0 ◆231 55 ◆
Claude Opus 4.6's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Opus 4.6. ◆ = independently verified score.
Where each leadsby capability area · benchmark wins
AreaClaude Mythos PreviewWINSClaude Opus 4.6
Biggest gaps
Claude Mythos Preview pulls furthest ahead on
- Humanity's Last Exam64.7 vs 39.9
- SWE-bench Pro77.8 vs 53.4
- CharXiv (RQ)93.2 vs 77.4
Claude Opus 4.6 pulls furthest ahead on
No ratified-area lead of 3 points or more.
Every shared benchmark15 · grouped by area
Reasoning 2
Coding 3
Agentic 2
Multimodal 1
Other shared benchmarks 7
Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.