Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Mythos Preview vs Claude Sonnet 4.6

14 SHARED BENCHMARKS

Across 14 shared benchmarks, Claude Mythos Preview scores higher on 14 and Claude Sonnet 4.6 on 0. The widest gap is Humanity's Last Exam, where Claude Mythos Preview scores 64.7 against 33.6.

ANTHROPICVSANTHROPIC14 SHARED14–0 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicAnthropic
Released—
Price per 1M tokens input / output—$3.00 / $15.00
Cost of 1M in + 1M out—$18.00
Head-to-head of 14 shared benchmarks14 wins0 wins
Scores tracked independently verified25 0 ◆152 24 ◆

Claude Sonnet 4.6's release date per Artificial Analysis. Prices: Anthropic's own price page for Claude Sonnet 4.6. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Mythos PreviewWINSClaude Sonnet 4.6
Reasoning20Claude Mythos Preview leads 2 of 2 · widest: Humanity's Last Exam 64.7 vs 33.6
Coding20Claude Mythos Preview leads 2 of 2 · widest: SWE-bench Pro 77.8 vs 58.1
Agentic20Claude Mythos Preview leads 2 of 2 · widest: BrowseComp 86.9 vs 74.7
Multimodal10Claude Mythos Preview leads 1 of 1 · widest: CharXiv (RQ) 93.2 vs 71.6

Biggest gaps

Claude Mythos Preview pulls furthest ahead on

  1. Humanity's Last Exam64.7 vs 33.6
  2. SWE-bench Pro77.8 vs 58.1
  3. CharXiv (RQ)93.2 vs 71.6

Claude Sonnet 4.6 pulls furthest ahead on

No ratified-area lead of 3 points or more.

Every shared benchmark14 · grouped by area

Multimodal 1

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.

Compare Claude Sonnet 4.6 withALL PAIRINGS →