Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Claude Sonnet 4.6 vs GPT-5.6 Luna

42 SHARED BENCHMARKS

Across 42 shared benchmarks, Claude Sonnet 4.6 scores higher on 9 and GPT-5.6 Luna on 33. The widest gap is AA-Omniscience · Non-hallucination, where Claude Sonnet 4.6 scores 51.6 against 7.4. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $3.00 per million input tokens).

ANTHROPICVSOPENAI42 SHARED9–33 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerAnthropicOpenAI
Released
Price per 1M tokens input / output$3.00 / $15.00$0.20 / $1.20
Cost of 1M in + 1M out$18.00$1.40 12.9× less
Head-to-head of 42 shared benchmarks9 wins33 wins
Scores tracked independently verified152 24 ◆201 43 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for Claude Sonnet 4.6; OpenAI's own price page for GPT-5.6 Luna. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaClaude Sonnet 4.6WINSGPT-5.6 Luna
Reasoning05GPT-5.6 Luna leads 5 of 5 · widest: CritPt 3.1 vs 20.6
Coding16GPT-5.6 Luna leads 6 of 7 · widest: SWE-bench Verified 79.6 vs 93
Agentic16GPT-5.6 Luna leads 6 of 7 · widest: Terminal-Bench 4.0 3 vs 11.6
Factuality12GPT-5.6 Luna leads 2 of 3 · widest: SimpleQA Verified 32.8 vs 41
Instruction Following10Claude Sonnet 4.6 leads 1 of 1 · widest: LiveBench · Instruction Following 63.2 vs 60.1
Long Context11Even, 1–1 of 2
Math01GPT-5.6 Luna leads 1 of 1
Multimodal12GPT-5.6 Luna leads 2 of 3 · widest: CharXiv (RQ) 71.6 vs 82.7

Biggest gaps

Claude Sonnet 4.6 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination51.6 vs 7.4
  2. MRCR v2 8 needle 128k (average)84.9 vs 74.8
  3. τ-Bench V3 · Banking34.4 vs 31.1

GPT-5.6 Luna pulls furthest ahead on

  1. CritPt20.6 vs 3.1
  2. Terminal-Bench 4.011.6 vs 3
  3. GDPVal47.2 vs 36

Every shared benchmark42 · grouped by area

Reasoning 5

ARC-AGI-258.3◆ 59.6
CritPt3.120.6
GPQA Diamond87.591.1
LiveBench · Reasoning84.8 ◆◆ 85.6

Coding 7

Agentic 7

BrowseComp74.783.3
GDPVal3647.2
Terminal-Bench 4.0311.6

Instruction Following 1

Multimodal 3

LMArena · Vision1283 ◆◆ 1259
CharXiv (RQ)71.682.7
MMMU-Pro73.378.6

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, openai-official), otherwise the lowest tracked offer.

Compare Claude Sonnet 4.6 withALL PAIRINGS →

Compare GPT-5.6 Luna withALL PAIRINGS →