Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-6 Luna vs Claude Opus 5

25 SHARED BENCHMARKS

Across 25 shared benchmarks, GPT-6 Luna scores higher on 2 and Claude Opus 5 on 23. The widest gap is Terminal-Bench 4.0, where Claude Opus 5 scores 49 against 12.6. GPT-6 Luna is the cheaper of the two on tracked API pricing ($0.10 against $5.00 per million input tokens).

OPENAIVSANTHROPIC25 SHARED2–23 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIAnthropic
Released
Price per 1M tokens input / output$0.10 / $0.50$5.00 / $25.00
Cost of 1M in + 1M out$0.60 50.0× less$30.00
Head-to-head of 25 shared benchmarks2 wins23 wins
Scores tracked independently verified102 33 ◆174 22 ◆

Release dates: Artificial Analysis for GPT-6 Luna; the vendor's own announcement for Claude Opus 5. Prices: OpenAI's own price page for GPT-6 Luna; Artificial Analysis for Claude Opus 5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-6 LunaWINSClaude Opus 5
Reasoning05Claude Opus 5 leads 5 of 5 · widest: ARC-AGI-2 59.3 vs 90.4
Coding04Claude Opus 5 leads 4 of 4 · widest: LiveBench · Agentic Coding 51.2 vs 65.2
Agentic02Claude Opus 5 leads 2 of 2 · widest: Terminal-Bench 4.0 12.6 vs 49
Factuality02Claude Opus 5 leads 2 of 2 · widest: AA-Omniscience · Non-hallucination 23.3 vs 39.2
Instruction Following01Claude Opus 5 leads 1 of 1 · widest: LiveBench · Instruction Following 55.9 vs 63.8
Long Context10GPT-6 Luna leads 1 of 1 · widest: AA-LCR 83.3 vs 79.3
Math01Claude Opus 5 leads 1 of 1 · widest: LiveBench · Mathematics 89.1 vs 95.7
Multimodal01Claude Opus 5 leads 1 of 1 · widest: MMMU-Pro 75.5 vs 84.7

Biggest gaps

GPT-6 Luna pulls furthest ahead on

  1. AA-LCR83.3 vs 79.3

Claude Opus 5 pulls furthest ahead on

  1. Terminal-Bench 4.049 vs 12.6
  2. AA-Omniscience · Non-hallucination39.2 vs 23.3
  3. ARC-AGI-290.4 vs 59.3

Every shared benchmark25 · grouped by area

Reasoning 5

ARC-AGI-30.1 ◆30.2
ARC-AGI-259.3 ◆90.4
CritPt19.429.1
LiveBench · Reasoning81.8 ◆◆ 91.2

Coding 4

LMArena · WebDev1593 ◆◆ 1693
LiveBench · Coding79 ◆◆ 81.5
SciCode54.656.4

Agentic 2

GDPVal43.460.4
Terminal-Bench 4.012.649

Instruction Following 1

Long Context 1

AA-LCR83.379.3

Math 1

Multimodal 1

MMMU-Pro75.584.7

Other shared benchmarks 8

ARC-AGI-186.7 ◆97.5
DeepSWE 1.166.668.8
livebench_data_analysis73.4 ◆◆ 74.5
livebench_language73.8 ◆◆ 88.7
OSWorld 2.052.770.6

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (openai-official, direct), otherwise the lowest tracked offer.

Compare Claude Opus 5 withALL PAIRINGS →