Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Mythos 5 vs GPT-5.6 Sol

Anthropic · released

Wins 3 of 4 areas

Coding · Agents · Images and charts

OpenAI · released

Wins 1 of 4 areas

Reasoning

Claude Mythos 5 is the stronger all-rounder.GPT-5.6 Sol is cheaper and better at reasoning.

Scores updated · 21 tests both models report · How we compare

Where each one wins

Tests won in each of the four areas where both have results. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.

Agents and images and charts rest on a single test each.

What it costs

Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.

Claude Mythos 5$10.00 to read · $50.00 to write$60Price from Anthropic's own price page
GPT-5.6 Sol$4.00 to read · $20.00 to write$24Price from Artificial Analysis

GPT-5.6 Sol costs 60% less for the same work.

The biggest differences

The tests each model wins by the widest margin, up to three each. Scores are out of 100.

Where Claude Mythos 5 pulls ahead

  • Long, multi-file coding tasks in real codebasesSWE-bench Pro+15.4points aheadClaude Mythos 580GPT-5.6 Sol64.6
  • Very hard expert questions across many subjectsHumanity's Last Exam+8.3points aheadClaude Mythos 557.8GPT-5.6 Sol49.5
  • Reasoning about charts from research papersCharXiv (RQ)+3.1points aheadClaude Mythos 588.9GPT-5.6 Sol85.8

Where GPT-5.6 Sol pulls ahead

  • Research-level physics problemsCritPt+3.7points aheadGPT-5.6 Sol32.3Claude Mythos 528.6
  • Abstract visual puzzles that people can solveARC-AGI-2+3.3points aheadGPT-5.6 Sol92.5Claude Mythos 589.2

Every test, side by side

All 21 tests both models report. The winning score is in its model's colour; marks a score checked independently.

CodingClaude Mythos 5

Full coding ranking

AgentsClaude Mythos 5

Full agents ranking

ReasoningGPT-5.6 Sol

Full reasoning ranking

Images and chartsClaude Mythos 5

Full images and charts ranking

Other results14 tests, not counted

Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.

Questions people ask

Which is better, Claude Mythos 5 or GPT-5.6 Sol?

Claude Mythos 5 wins three of the four areas where both have results: coding, agents and images and charts. GPT-5.6 Sol wins reasoning, and costs 60% less.

Which is better for coding?

Claude Mythos 5. It wins 1 of the 2 coding tests both models report; GPT-5.6 Sol wins none, and 1 is a tie.

Which is cheaper?

Claude Mythos 5 costs $10.00 per million input tokens and $50.00 per million output tokens; GPT-5.6 Sol costs $4.00 and $20.00. That makes GPT-5.6 Sol about 60% cheaper for the same work.

How do you compare the two?

We use the 21 benchmark tests both models have published scores on. The verdict counts the 7 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 14 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.