Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Luna vs Qwen3.5 397B A17B

34 SHARED BENCHMARKS

Across 34 shared benchmarks, GPT-5.6 Luna scores higher on 32 and Qwen3.5 397B A17B on 2. The widest gap is AA Agentic Index, where GPT-5.6 Luna scores 42.7 against 10.6. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $0.60 per million input tokens).

OPENAIVSALIBABA34 SHARED32–2 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIAlibaba
Released
Price per 1M tokens input / output$0.20 / $1.20$0.60 / $3.60
Cost of 1M in + 1M out$1.40 3.0× less$4.20
Head-to-head of 34 shared benchmarks32 wins2 wins
Scores tracked independently verified201 43 ◆162 11 ◆

Release dates per Artificial Analysis. Prices: Artificial Analysis for GPT-5.6 Luna; Artificial Analysis for Qwen3.5 397B A17B. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.6 LunaWINSQwen3.5 397B A17B
Reasoning30GPT-5.6 Luna leads 3 of 3 · widest: Humanity's Last Exam 39.5 vs 29
Coding50GPT-5.6 Luna leads 5 of 5 · widest: Terminal-Bench 2.1 80.9 vs 51.3
Agentic60GPT-5.6 Luna leads 6 of 6 · widest: GDPVal 47.2 vs 14
Factuality21GPT-5.6 Luna leads 2 of 3 · widest: SimpleQA Verified 41 vs 26
Long Context10GPT-5.6 Luna leads 1 of 1 · widest: AA-LCR 83.7 vs 77.3
Math30GPT-5.6 Luna leads 3 of 3 · widest: FrontierMath Tiers 1-3 (v2) 82.1 vs 31.2
Multimodal11Even, 1–1 of 2

Biggest gaps

GPT-5.6 Luna pulls furthest ahead on

  1. GDPVal47.2 vs 14
  2. FrontierMath Tiers 1-3 (v2)82.1 vs 31.2
  3. AA ApexAgents35.8 vs 15.3

Qwen3.5 397B A17B pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination11.1 vs 7.4

Every shared benchmark34 · grouped by area

Reasoning 3

Coding 5

Agentic 6

GDPVal47.214
Terminal-Bench 4.011.60

Long Context 1

Math 3

Multimodal 2

LMArena · Vision1259 ◆◆ 1262
MMMU-Pro78.677.3

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (direct, direct), otherwise the lowest tracked offer.

Compare GPT-5.6 Luna withALL PAIRINGS →

Compare Qwen3.5 397B A17B withALL PAIRINGS →