Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Fable 5.1 vs GPT-6 Sol

25 SHARED BENCHMARKS

Across 25 shared benchmarks, Claude Fable 5.1 scores higher on 21 and GPT-6 Sol on 4. The widest gap is AA-Omniscience, where Claude Fable 5.1 scores 43.5 against 27.1. GPT-6 Sol is the cheaper of the two on tracked API pricing ($2.00 against $10.00 per million input tokens).

ANTHROPICVSOPENAI25 SHARED21–4 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicOpenAI
Released
Price per 1M tokens input / output$10.00 / $50.00$2.00 / $10.00
Cost of 1M in + 1M out$60.00$12.00 5.0× less
Head-to-head of 25 shared benchmarks21 wins4 wins
Scores tracked independently verified144 25 ◆107 37 ◆

Release dates: the vendor's own announcement for Claude Fable 5.1; Artificial Analysis for GPT-6 Sol. Prices: Anthropic's own price page for Claude Fable 5.1; Artificial Analysis for GPT-6 Sol. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Fable 5.1WINSGPT-6 Sol
Reasoning41Claude Fable 5.1 leads 4 of 5 · widest: Humanity's Last Exam 59.1 vs 47.9
Coding40Claude Fable 5.1 leads 4 of 4 · widest: LiveBench · Agentic Coding 66.1 vs 52.9
Agentic30Claude Fable 5.1 leads 3 of 3 · widest: GDPVal 61.7 vs 49.3
Factuality11Even, 1–1 of 2
Instruction Following10Claude Fable 5.1 leads 1 of 1 · widest: LiveBench · Instruction Following 73 vs 68.6
Long Context10Claude Fable 5.1 leads 1 of 1
Math10Claude Fable 5.1 leads 1 of 1

Biggest gaps

Claude Fable 5.1 pulls furthest ahead on

  1. GDPVal61.7 vs 49.3
  2. LiveBench · Agentic Coding66.1 vs 52.9
  3. Humanity's Last Exam59.1 vs 47.9

GPT-6 Sol pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination39.9 vs 27.4

Every shared benchmark25 · grouped by area

Reasoning 5

ARC-AGI-290◆ 89.6
CritPt29.730.9
LiveBench · Reasoning91.7 ◆◆ 88.7
SimpleBench86.6 ◆◆ 73.1

Coding 4

LMArena · WebDev1751 ◆◆ 1681
LiveBench · Coding86.4 ◆◆ 81.8
SciCode63.157.6

Agentic 3

GDPVal61.749.3
Terminal-Bench 4.05243.9

Instruction Following 1

Long Context 1

AA-LCR85.383.7

Math 1

Other shared benchmarks 8

ARC-AGI-197.5◆ 95.5
DeepSWE 1.167.468.8
GDPval-AA 2.117351487
livebench_data_analysis80.3 ◆◆ 81.2
livebench_language89.5 ◆◆ 85.3
OSWorld 2.077.964.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (anthropic-official, direct), otherwise the lowest tracked offer.

Compare Claude Fable 5.1 withALL PAIRINGS →