Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.5 vs GPT-5.6 Luna

69 SHARED BENCHMARKS

Across 69 shared benchmarks, GPT-5.5 scores higher on 46 and GPT-5.6 Luna on 22, with 1 level. The widest gap is ARC-AGI-3, where GPT-5.5 scores 0.4 against 0.2. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $5.00 per million input tokens).

OPENAIVSOPENAI69 SHARED46–22 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIOpenAI
Released
Price per 1M tokens input / output$5.00 / $30.00$0.20 / $1.20
Cost of 1M in + 1M out$35.00$1.40 25.0× less
Head-to-head of 69 shared benchmarks46 wins22 wins
Scores tracked independently verified185 29 ◆201 43 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for GPT-5.5; Artificial Analysis for GPT-5.6 Luna. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.5WINSGPT-5.6 Luna
Reasoning70GPT-5.5 leads 7 of 7 · widest: SimpleBench 69 vs 46.8
Coding33Even, 3–3 of 6
Agentic61GPT-5.5 leads 6 of 7 · widest: Terminal-Bench 4.0 14.6 vs 11.6
Factuality30GPT-5.5 leads 3 of 3 · widest: SimpleQA Verified 63 vs 41
Instruction Following10GPT-5.5 leads 1 of 1 · widest: LiveBench · Instruction Following 70.7 vs 60.1
Long Context10GPT-5.5 leads 1 of 1
Math41GPT-5.5 leads 4 of 5 · widest: FrontierMath Tier 4 72.5 vs 58.5
Multimodal20GPT-5.5 leads 2 of 2 · widest: LMArena · Vision 1292 vs 1259

Biggest gaps

GPT-5.5 pulls furthest ahead on

  1. SimpleQA Verified63 vs 41
  2. AA-Omniscience · Non-hallucination11 vs 7.4
  3. SimpleBench69 vs 46.8

GPT-5.6 Luna pulls furthest ahead on

  1. GDPVal47.2 vs 41.8
  2. SWE-bench Pro62.7 vs 58.6
  3. LMArena · WebDev1520 vs 1512

Every shared benchmark69 · grouped by area

Reasoning 7

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
ARC-AGI-30.4 ◆0.2
ARC-AGI-285◆ 59.6
CritPt27.120.6
GPQA Diamond93.591.1
LiveBench · Reasoning89.7 ◆◆ 85.6
SimpleBench69 ◆◆ 46.8

Coding 6

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
LMArena · WebDev1512 ◆◆ 1520
LiveBench · Coding82.2 ◆◆ 82.9
SciCode55.853.6

Agentic 7

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
BrowseComp84.483.3
GDPVal41.847.2
Terminal-Bench 4.014.611.6

Instruction Following 1

Long Context 1

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
AA-LCR84.383.7

Math 5

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
AIME 2026100 ◆97.6
HMMT Feb. 202698.5 ◆98.5
LiveBench · Mathematics95.9 ◆◆ 87.2

Multimodal 2

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
LMArena · Vision1292 ◆◆ 1259
MMMU-Pro79.978.6

Other shared benchmarks 37

BenchmarkGPT-5.5MARGINGPT-5.6 Luna
ARC-AGI-195 ◆◆ 90.7
DeepSWE7067.2
Dynamic Benchmarks - Emotional reliance0.91
Dynamic Benchmarks - Mental health0.81
Dynamic Benchmarks - Self-harm0.90.9
GDPval-AA v215091530
HealthBench56.555.8
Image input evaluations - extremism11
Image input evaluations - harms-erotic11
Image input evaluations - hate11
Image input evaluations - self-harm11
livebench_language87.4 ◆◆ 72.6
Production Benchmarks - Gore0.80.6
Toolathlon55.653.4
U18 evaluations - Age-restricted goods, services, and dangerous challenges / activities0.70.8
U18 evaluations - Eating Disorders0.60.7
U18 evaluations - Emotional Reliance0.90.9
U18 evaluations - Gore0.80.8
U18 evaluations - Self Harm11
U18 evaluations - Sexual Content0.90.9

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare GPT-5.5 withALL PAIRINGS →

Compare GPT-5.6 Luna withALL PAIRINGS →