Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Luna vs GPT-5.6 Sol

57 SHARED BENCHMARKS

Across 57 shared benchmarks, GPT-5.6 Luna scores higher on 0 and GPT-5.6 Sol on 57. The widest gap is Terminal-Bench 4.0, where GPT-5.6 Sol scores 39.9 against 11.6. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $4.00 per million input tokens).

OPENAIVSOPENAI57 SHARED0–57 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIOpenAI
Released
Price per 1M tokens input / output$0.20 / $1.20$4.00 / $20.00
Cost of 1M in + 1M out$1.40 17.1× less$24.00
Head-to-head of 57 shared benchmarks0 wins57 wins
Scores tracked independently verified183 43 ◆218 31 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for GPT-5.6 Luna; Artificial Analysis for GPT-5.6 Sol. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.6 LunaWINSGPT-5.6 Sol
Reasoning07GPT-5.6 Sol leads 7 of 7 · widest: CritPt 20.6 vs 32.3
Coding06GPT-5.6 Sol leads 6 of 6 · widest: LiveBench · Agentic Coding 48.4 vs 56.2
Agentic07GPT-5.6 Sol leads 7 of 7 · widest: Terminal-Bench 4.0 11.6 vs 39.9
Factuality03GPT-5.6 Sol leads 3 of 3 · widest: SimpleQA Verified 41 vs 69.7
Instruction Following01GPT-5.6 Sol leads 1 of 1 · widest: LiveBench · Instruction Following 60.1 vs 71.8
Long Context01GPT-5.6 Sol leads 1 of 1
Math04GPT-5.6 Sol leads 4 of 4 · widest: FrontierMath Tier 4 58.5 vs 83
Multimodal03GPT-5.6 Sol leads 3 of 3 · widest: MMMU-Pro 78.6 vs 83.4

Biggest gaps

GPT-5.6 Luna pulls furthest ahead on

No ratified-area lead of 3 points or more.

GPT-5.6 Sol pulls furthest ahead on

  1. Terminal-Bench 4.039.9 vs 11.6
  2. SimpleQA Verified69.7 vs 41
  3. CritPt32.3 vs 20.6

Every shared benchmark57 · grouped by area

Reasoning 7

ARC-AGI-30.27.8
ARC-AGI-259.6 ◆◆ 92.5
CritPt20.632.3
GPQA Diamond91.194.1
LiveBench · Reasoning85.6 ◆◆ 91.7
SimpleBench46.8 ◆◆ 64.8

Coding 6

LMArena · WebDev1520 ◆◆ 1617
LiveBench · Coding82.9 ◆◆ 83.9
SciCode53.657.1

Agentic 7

BrowseComp83.390.4
GDPVal47.254.4
Terminal-Bench 4.011.639.9

Instruction Following 1

Long Context 1

AA-LCR83.784

Multimodal 3

LMArena · Vision1259 ◆◆ 1281
CharXiv (RQ)82.785.8
MMMU-Pro78.683.4

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare GPT-5.6 Sol withALL PAIRINGS →