Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Sonnet 5 vs GPT-5.6 Luna

45 SHARED BENCHMARKS

Across 45 shared benchmarks, Claude Sonnet 5 scores higher on 30 and GPT-5.6 Luna on 13, with 2 level. The widest gap is AA-Omniscience · Non-hallucination, where Claude Sonnet 5 scores 60.6 against 7.4. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $2.00 per million input tokens).

ANTHROPICVSOPENAI45 SHARED30–13 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicOpenAI
Released
Price per 1M tokens input / output$2.00 / $10.00$0.20 / $1.20
Cost of 1M in + 1M out$12.00$1.40 8.6× less
Head-to-head of 45 shared benchmarks30 wins13 wins
Scores tracked independently verified153 20 ◆201 43 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for Claude Sonnet 5; OpenAI's own price page for GPT-5.6 Luna. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Sonnet 5WINSGPT-5.6 Luna
Reasoning31Claude Sonnet 5 leads 3 of 5 · widest: SimpleBench 60.6 vs 46.8
Coding43Claude Sonnet 5 leads 4 of 7 · widest: LiveBench · Agentic Coding 59.4 vs 48.4
Agentic50Claude Sonnet 5 leads 5 of 5 · widest: τ-Bench V3 · Banking 37.3 vs 31.1
Factuality12GPT-5.6 Luna leads 2 of 3 · widest: SimpleQA Verified 33.7 vs 41
Instruction Following10Claude Sonnet 5 leads 1 of 1 · widest: LiveBench · Instruction Following 63.9 vs 60.1
Long Context11Even, 1–1 of 2
Math12GPT-5.6 Luna leads 2 of 3 · widest: FrontierMath Tier 4 29.3 vs 58.5
Multimodal21Claude Sonnet 5 leads 2 of 3 · widest: CharXiv (RQ) 88.3 vs 82.7

Biggest gaps

Claude Sonnet 5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination60.6 vs 7.4
  2. SimpleBench60.6 vs 46.8
  3. LiveBench · Agentic Coding59.4 vs 48.4

GPT-5.6 Luna pulls furthest ahead on

  1. FrontierMath Tier 458.5 vs 29.3
  2. FrontierMath Tiers 1-3 (v2)82.1 vs 65.6
  3. CritPt20.6 vs 16.9

Every shared benchmark45 · grouped by area

Reasoning 5

CritPt16.920.6
GPQA Diamond91.191.1
LiveBench · Reasoning88.7 ◆◆ 85.6
SimpleBench60.6 ◆◆ 46.8

Coding 7

Agentic 5

BrowseComp84.783.3
GDPVal47.547.2
Terminal-Bench 4.014.111.6

Instruction Following 1

Long Context 2

Math 3

Multimodal 3

LMArena · Vision1275 ◆◆ 1259
CharXiv (RQ)88.382.7
MMMU-Pro77.378.6

Other shared benchmarks 16

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, openai-official), otherwise the lowest tracked offer.

Compare Claude Sonnet 5 withALL PAIRINGS →

Compare GPT-5.6 Luna withALL PAIRINGS →