Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.4 vs GPT-5.6 Luna

60 SHARED BENCHMARKS

Across 60 shared benchmarks, GPT-5.4 scores higher on 36 and GPT-5.6 Luna on 24. The widest gap is FrontierMath (overall), where GPT-5.6 Luna scores 78.6 against 47.6. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $2.50 per million input tokens).

OPENAIVSOPENAI60 SHARED36–24 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIOpenAI
Released
Price per 1M tokens input / output$2.50 / $15.00$0.20 / $1.20
Cost of 1M in + 1M out$17.50$1.40 12.5× less
Head-to-head of 60 shared benchmarks36 wins24 wins
Scores tracked independently verified199 41 ◆201 43 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for GPT-5.4; Artificial Analysis for GPT-5.6 Luna. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.4WINSGPT-5.6 Luna
Reasoning60GPT-5.4 leads 6 of 6 · widest: ARC-AGI-2 73.3 vs 59.6
Coding24GPT-5.6 Luna leads 4 of 6 · widest: SWE-bench Pro 57.7 vs 62.7
Agentic24GPT-5.6 Luna leads 4 of 6 · widest: GDPVal 36.6 vs 47.2
Factuality30GPT-5.4 leads 3 of 3 · widest: AA-Omniscience · Accuracy 50.9 vs 42.7
Instruction Following10GPT-5.4 leads 1 of 1 · widest: LiveBench · Instruction Following 70.2 vs 60.1
Long Context01GPT-5.6 Luna leads 1 of 1
Math23GPT-5.6 Luna leads 3 of 5 · widest: FrontierMath Tier 4 49 vs 58.5
Multimodal21GPT-5.4 leads 2 of 3 · widest: LMArena · Vision 1302 vs 1259

Biggest gaps

GPT-5.4 pulls furthest ahead on

  1. τ-Bench V3 · Banking39.6 vs 31.1
  2. ARC-AGI-273.3 vs 59.6
  3. AA-Omniscience · Accuracy50.9 vs 42.7

GPT-5.6 Luna pulls furthest ahead on

  1. GDPVal47.2 vs 36.6
  2. FrontierMath Tier 458.5 vs 49
  3. AA IT-Bench SRE40.3 vs 34.5

Every shared benchmark60 · grouped by area

Reasoning 6

BenchmarkGPT-5.4MARGINGPT-5.6 Luna
ARC-AGI-30.2 ◆0.2
ARC-AGI-273.3◆ 59.6
CritPt23.420.6
LiveBench · Reasoning88.1 ◆◆ 85.6

Coding 6

BenchmarkGPT-5.4MARGINGPT-5.6 Luna
LMArena · WebDev1466 ◆◆ 1520
LiveBench · Coding77.5 ◆◆ 82.9
SciCode56.653.6

Agentic 6

Instruction Following 1

Long Context 1

BenchmarkGPT-5.4MARGINGPT-5.6 Luna
AA-LCR8283.7

Math 5

BenchmarkGPT-5.4MARGINGPT-5.6 Luna
AIME 202699.2 ◆97.6
HMMT Feb. 202697.7 ◆98.5
LiveBench · Mathematics94.2 ◆◆ 87.2

Multimodal 3

BenchmarkGPT-5.4MARGINGPT-5.6 Luna
LMArena · Vision1302 ◆◆ 1259
CharXiv (RQ)82.882.7
MMMU-Pro78.478.6

Other shared benchmarks 29

BenchmarkGPT-5.4MARGINGPT-5.6 Luna
ARC-AGI-193.7 ◆◆ 90.7
Dynamic Benchmarks - Emotional reliance11
Dynamic Benchmarks - Mental health0.91
Dynamic Benchmarks - Self-harm10.9
Image input evaluations - extremism11
Image input evaluations - harms-erotic11
Image input evaluations - hate11
Image input evaluations - self-harm11
livebench_language82.6 ◆◆ 72.6
Production Benchmarks - Gore0.80.6
Toolathlon54.653.4
U18 evaluations - Age-restricted goods, services, and dangerous challenges / activities0.80.8
U18 evaluations - Eating Disorders0.70.7
U18 evaluations - Emotional Reliance0.90.9
U18 evaluations - Gore0.80.8
U18 evaluations - Self Harm11
U18 evaluations - Sexual Content0.90.9

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare GPT-5.4 withALL PAIRINGS →

Compare GPT-5.6 Luna withALL PAIRINGS →