Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Luna vs Kimi K2.5

33 SHARED BENCHMARKS

Across 33 shared benchmarks, GPT-5.6 Luna scores higher on 29 and Kimi K2.5 on 4. The widest gap is CritPt, where GPT-5.6 Luna scores 20.6 against 3.1. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $0.60 per million input tokens).

OPENAIVSMOONSHOT33 SHARED29–4 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIMoonshot
Released
Price per 1M tokens input / output$0.20 / $1.20$0.60 / $3.00
Cost of 1M in + 1M out$1.40 2.6× less$3.60
Head-to-head of 33 shared benchmarks29 wins4 wins
Scores tracked independently verified201 43 ◆172 26 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for GPT-5.6 Luna; Artificial Analysis for Kimi K2.5. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.6 LunaWINSKimi K2.5
Reasoning40GPT-5.6 Luna leads 4 of 4 · widest: CritPt 20.6 vs 3.1
Coding50GPT-5.6 Luna leads 5 of 5 · widest: Terminal-Bench 2.1 80.9 vs 45.7
Agentic50GPT-5.6 Luna leads 5 of 5 · widest: AA ApexAgents 35.8 vs 11.5
Factuality21GPT-5.6 Luna leads 2 of 3 · widest: AA-Omniscience · Accuracy 42.7 vs 35.2
Long Context10GPT-5.6 Luna leads 1 of 1 · widest: AA-LCR 83.7 vs 78
Math20GPT-5.6 Luna leads 2 of 2 · widest: HMMT Feb. 2026 98.5 vs 87.1
Multimodal21GPT-5.6 Luna leads 2 of 3 · widest: CharXiv (RQ) 82.7 vs 77.5

Biggest gaps

GPT-5.6 Luna pulls furthest ahead on

  1. CritPt20.6 vs 3.1
  2. ARC-AGI-259.6 vs 11.8
  3. AA ApexAgents35.8 vs 11.5

Kimi K2.5 pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination34.3 vs 7.4
  2. LMArena · Vision1269 vs 1259

Every shared benchmark33 · grouped by area

Reasoning 4

BenchmarkGPT-5.6 LunaMARGINKimi K2.5
ARC-AGI-259.6 ◆◆ 11.8
CritPt20.63.1
GPQA Diamond91.187.9

Coding 5

Agentic 5

Long Context 1

BenchmarkGPT-5.6 LunaMARGINKimi K2.5
AA-LCR83.778

Math 2

BenchmarkGPT-5.6 LunaMARGINKimi K2.5
AIME 202697.6◆ 95.8
HMMT Feb. 202698.5◆ 87.1

Multimodal 3

BenchmarkGPT-5.6 LunaMARGINKimi K2.5
LMArena · Vision1259 ◆◆ 1269
CharXiv (RQ)82.777.5
MMMU-Pro78.675.4

Other shared benchmarks 10

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare GPT-5.6 Luna withALL PAIRINGS →

Compare Kimi K2.5 withALL PAIRINGS →