Claude Mythos Preview vs Gemini 3.1 Pro
13 SHARED BENCHMARKSAcross 13 shared benchmarks, Claude Mythos Preview scores higher on 13 and Gemini 3.1 Pro on 0. The widest gap is CyberGym, where Claude Mythos Preview scores 83.1 against 38.8.
ANTHROPICVSGOOGLE13 SHARED13–0 HEAD-TO-HEADNEWEST SCORE ADDED
At a glance
MakerAnthropicGoogle
Released—
Price per 1M tokens input / output—$2.00 / $12.00
Cost of 1M in + 1M out—$14.00
Head-to-head of 13 shared benchmarks13 wins0 wins
Scores tracked independently verified25 0 ◆215 30 ◆
Gemini 3.1 Pro's release date per Artificial Analysis. Prices: Artificial Analysis for Gemini 3.1 Pro. ◆ = independently verified score.
Where each leadsby capability area · benchmark wins
AreaClaude Mythos PreviewWINSGemini 3.1 Pro
Biggest gaps
Claude Mythos Preview pulls furthest ahead on
- SWE-bench Pro77.8 vs 54.2
- Humanity's Last Exam64.7 vs 47
- SWE-bench Verified93.9 vs 80.6
Gemini 3.1 Pro pulls furthest ahead on
No ratified-area lead of 3 points or more.
Every shared benchmark13 · grouped by area
Reasoning 2
Coding 3
Agentic 2
Multimodal 1
Other shared benchmarks 5
Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing.