Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Mythos Preview vs Claude Opus 4.8

22 SHARED BENCHMARKS

Across 22 shared benchmarks, Claude Mythos Preview scores higher on 20 and Claude Opus 4.8 on 2. The widest gap is SWE-bench Multimodal, where Claude Mythos Preview scores 59 against 38.4.

ANTHROPICVSANTHROPIC22 SHARED20–2 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicAnthropic
Released—
Price per 1M tokens input / output—$5.00 / $25.00
Cost of 1M in + 1M out—$30.00
Head-to-head of 22 shared benchmarks20 wins2 wins
Scores tracked independently verified25 0 ◆200 32 ◆

Claude Opus 4.8's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Opus 4.8. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Mythos PreviewWINSClaude Opus 4.8
Reasoning20Claude Mythos Preview leads 2 of 2 · widest: Humanity's Last Exam 64.7 vs 48.7
Coding30Claude Mythos Preview leads 3 of 3 · widest: SWE-bench Pro 77.8 vs 69.2
Agentic11Even, 1–1 of 2
Multimodal10Claude Mythos Preview leads 1 of 1 · widest: CharXiv (RQ) 93.2 vs 89.9

Biggest gaps

Claude Mythos Preview pulls furthest ahead on

  1. Humanity's Last Exam64.7 vs 48.7
  2. SWE-bench Pro77.8 vs 69.2
  3. SWE-bench Verified93.9 vs 88.6

Claude Opus 4.8 pulls furthest ahead on

  1. OSWorld-Verified83.4 vs 79.6

Every shared benchmark22 · grouped by area

Multimodal 1

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Claude Opus 4.8 withALL PAIRINGS →