Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Luna vs Claude Opus 5

49 SHARED BENCHMARKS

Across 49 shared benchmarks, GPT-5.6 Luna scores higher on 4 and Claude Opus 5 on 45. The widest gap is AA-Omniscience · Non-hallucination, where Claude Opus 5 scores 39.2 against 7.4. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $5.00 per million input tokens).

OPENAIVSANTHROPIC49 SHARED4–45 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIAnthropic
Released
Price per 1M tokens input / output$0.20 / $1.20$5.00 / $25.00
Cost of 1M in + 1M out$1.40 21.4× less$30.00
Head-to-head of 49 shared benchmarks4 wins45 wins
Scores tracked independently verified183 43 ◆174 22 ◆

Release dates: Artificial Analysis for GPT-5.6 Luna; the vendor's own announcement for Claude Opus 5. Prices: OpenAI's own price page for GPT-5.6 Luna; Artificial Analysis for Claude Opus 5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.6 LunaWINSClaude Opus 5
Reasoning07Claude Opus 5 leads 7 of 7 · widest: SimpleBench 46.8 vs 80.6
Coding16Claude Opus 5 leads 6 of 7 · widest: LiveBench · Agentic Coding 48.4 vs 65.2
Agentic05Claude Opus 5 leads 5 of 5 · widest: Terminal-Bench 4.0 11.6 vs 49
Factuality03Claude Opus 5 leads 3 of 3 · widest: AA-Omniscience · Non-hallucination 7.4 vs 39.2
Instruction Following01Claude Opus 5 leads 1 of 1 · widest: LiveBench · Instruction Following 60.1 vs 63.8
Long Context10GPT-5.6 Luna leads 1 of 1 · widest: AA-LCR 83.7 vs 79.3
Math03Claude Opus 5 leads 3 of 3 · widest: FrontierMath Tier 4 58.5 vs 73.2
Multimodal03Claude Opus 5 leads 3 of 3 · widest: MMMU-Pro 78.6 vs 84.7

Biggest gaps

GPT-5.6 Luna pulls furthest ahead on

  1. AA-LCR83.7 vs 79.3

Claude Opus 5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination39.2 vs 7.4
  2. Terminal-Bench 4.049 vs 11.6
  3. SimpleBench80.6 vs 46.8

Every shared benchmark49 · grouped by area

Reasoning 7

ARC-AGI-30.230.2
ARC-AGI-259.6 ◆90.4
CritPt20.629.1
GPQA Diamond91.193.2
LiveBench · Reasoning85.6 ◆◆ 91.2
SimpleBench46.8 ◆◆ 80.6

Coding 7

Agentic 5

BrowseComp83.390.8
GDPVal47.260.4
Terminal-Bench 4.011.649

Instruction Following 1

Long Context 1

AA-LCR83.779.3

Math 3

Multimodal 3

LMArena · Vision1259 ◆◆ 1321
CharXiv (RQ)82.783.7
MMMU-Pro78.684.7

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (openai-official, direct), otherwise the lowest tracked offer.

Compare Claude Opus 5 withALL PAIRINGS →