Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Luna vs Qwen3.8 Max Preview

38 SHARED BENCHMARKS

Across 38 shared benchmarks, GPT-5.6 Luna scores higher on 9 and Qwen3.8 Max Preview on 29. The widest gap is AA-Omniscience · Non-hallucination, where Qwen3.8 Max Preview scores 71.2 against 7.4. GPT-5.6 Luna is the cheaper of the two on tracked API pricing ($0.20 against $2.00 per million input tokens).

OPENAIVSALIBABA38 SHARED9–29 HEAD-TO-HEADNEWEST SCORE ADDED

At a glance

MakerOpenAIAlibaba
Released
Price per 1M tokens input / output$0.20 / $1.20$2.00 / $6.00
Cost of 1M in + 1M out$1.40 5.7× less$8.00
Head-to-head of 38 shared benchmarks9 wins29 wins
Scores tracked independently verified201 43 ◆85 19 ◆

Release dates per Artificial Analysis. Prices: OpenAI's own price page for GPT-5.6 Luna; Artificial Analysis for Qwen3.8 Max Preview. ◆ = independently verified score.

Where each leadsby capability area · benchmark wins

AreaGPT-5.6 LunaWINSQwen3.8 Max Preview
Reasoning13Qwen3.8 Max Preview leads 3 of 4 · widest: Humanity's Last Exam 39.5 vs 43.1
Coding24Qwen3.8 Max Preview leads 4 of 6 · widest: LiveBench · Agentic Coding 48.4 vs 64.7
Agentic04Qwen3.8 Max Preview leads 4 of 4 · widest: Terminal-Bench 4.0 11.6 vs 38.9
Factuality12Qwen3.8 Max Preview leads 2 of 3 · widest: AA-Omniscience · Non-hallucination 7.4 vs 71.2
Instruction Following01Qwen3.8 Max Preview leads 1 of 1 · widest: LiveBench · Instruction Following 60.1 vs 74.1
Long Context10GPT-5.6 Luna leads 1 of 1 · widest: AA-LCR 83.7 vs 80.3
Math21GPT-5.6 Luna leads 2 of 3 · widest: FrontierMath Tier 4 58.5 vs 46.3
Multimodal02Qwen3.8 Max Preview leads 2 of 2 · widest: MMMU-Pro 78.6 vs 82.8

Biggest gaps

GPT-5.6 Luna pulls furthest ahead on

  1. AA-Omniscience · Accuracy42.7 vs 31.7
  2. FrontierMath Tier 458.5 vs 46.3
  3. LiveBench · Coding82.9 vs 72.9

Qwen3.8 Max Preview pulls furthest ahead on

  1. AA-Omniscience · Non-hallucination71.2 vs 7.4
  2. Terminal-Bench 4.038.9 vs 11.6
  3. τ-Bench V3 · Banking47.8 vs 31.1

Every shared benchmark38 · grouped by area

Reasoning 4

Coding 6

Agentic 4

GDPVal47.258.4
Terminal-Bench 4.011.638.9

Instruction Following 1

Long Context 1

Multimodal 2

LMArena · Vision1259 ◆◆ 1314
MMMU-Pro78.682.8

Each score is the one the model's own page shows — the most authoritative tracked result for that benchmark, on the benchmark's standard methodology (source-attributed). ↓ marks lower-is-better metrics; ◆ an independently verified score. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (openai-official, direct), otherwise the lowest tracked offer.

Compare GPT-5.6 Luna withALL PAIRINGS →

Compare Qwen3.8 Max Preview withALL PAIRINGS →