Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Mythos Preview vs Claude Sonnet 5

14 SHARED BENCHMARKS

Across 14 shared benchmarks, Claude Mythos Preview scores higher on 13 and Claude Sonnet 5 on 1. The widest gap is SWE-bench Multimodal, where Claude Mythos Preview scores 59 against 28.1.

ANTHROPICVSANTHROPIC14 SHARED13–1 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicAnthropic
Released—
Price per 1M tokens input / output—$2.00 / $10.00
Cost of 1M in + 1M out—$12.00
Head-to-head of 14 shared benchmarks13 wins1 wins
Scores tracked independently verified25 0 ◆153 20 ◆

Claude Sonnet 5's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Sonnet 5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Mythos PreviewWINSClaude Sonnet 5
Reasoning20Claude Mythos Preview leads 2 of 2 · widest: Humanity's Last Exam 64.7 vs 41.3
Coding30Claude Mythos Preview leads 3 of 3 · widest: SWE-bench Pro 77.8 vs 63.2
Agentic11Even, 1–1 of 2
Multimodal10Claude Mythos Preview leads 1 of 1 · widest: CharXiv (RQ) 93.2 vs 88.3

Biggest gaps

Claude Mythos Preview pulls furthest ahead on

  1. Humanity's Last Exam64.7 vs 41.3
  2. SWE-bench Pro77.8 vs 63.2
  3. SWE-bench Multilingual87.3 vs 78.3

Claude Sonnet 5 pulls furthest ahead on

No ratified-area lead of 3 points or more.

Every shared benchmark14 · grouped by area

Multimodal 1

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Claude Sonnet 5 withALL PAIRINGS →