Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.6 Luna

Score basis

GPT-5.6 Luna is strong in long context; capable in reasoning, coding, agentic tasks, and multimodal tasks; and behind the leaders in math and factuality. Too few results yet to rate safety, multilingual tasks, or instruction following.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong1 capability
  1. Long Context−5.0%2 of 3
Capable4 capabilities
  1. Reasoning−16.2%6 of 6
  2. Coding−20.7%7 of 10
  3. Agentic−21.0%5 of 7
  4. Multimodal−22.8%2 of 6
Limited2 capabilities
  1. Math−26.9%3 of 5
  2. Factuality−36.7%3 of 4
Not rated3 capabilities

Too few results yet to rate Safety, Multilingual or Instruction Following.

Price

$0.20input$1.20outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Cheaper than 60% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

184results on65benchmarks

  • 43 independently verified
  • 96 aggregator
  • 22 vendor-reported
  • 23 cross-referenced

From 17 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

GPT-5.6 Luna benchmark results

184 results on 65 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

StrongFull long context ranking

5.0% behind the leader2 of 3 ranked benchmarks measured

Show 5 more long context resultsHide 5 long context results

Reasoning

CapableFull reasoning ranking

16.2% behind the leader6 of 6 ranked benchmarks measured

Show 34 more reasoning resultsHide 34 reasoning results

Coding

CapableFull coding ranking

20.7% behind the leader7 of 10 ranked benchmarks measured

Show 14 more coding resultsHide 14 coding results

Agentic

CapableFull agentic ranking

21.0% behind the leader5 of 7 ranked benchmarks measured

Show 17 more agentic resultsHide 17 agentic results

Multimodal

CapableFull multimodal ranking

22.8% behind the leader2 of 6 ranked benchmarks measured

Show 6 more multimodal resultsHide 6 multimodal results

26.9% behind the leader3 of 5 ranked benchmarks measured

Show 5 more math resultsHide 5 math results

Factuality

LimitedFull factuality ranking

36.7% behind the leader3 of 4 ranked benchmarks measured

Show 10 more factuality resultsHide 10 factuality results

Instruction Following

Not enough dataFull instruction following ranking

0 of 3 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 58 more resultsHide 58 results

GPT-5.6 Luna: common questions

Who makes GPT-5.6 Luna?

GPT-5.6 Luna is made by OpenAI.

When was GPT-5.6 Luna released?

GPT-5.6 Luna was released on Jul 9, 2026, according to Artificial Analysis.

What is GPT-5.6 Luna good at?

GPT-5.6 Luna is strong in long context; capable in reasoning, coding, agentic tasks, and multimodal tasks; and behind the leaders in math and factuality. Too few results yet to rate safety, multilingual tasks, or instruction following.

How much does GPT-5.6 Luna cost?

GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 60% of the 331 priced models we track.

How many benchmarks has GPT-5.6 Luna been tested on?

We track 184 results for GPT-5.6 Luna on 65 benchmarks from 17 sources, 43 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does GPT-5.6 Luna support?

OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.6 Luna.

About this record

Where GPT-5.6 Luna's numbers come from, and every name it appears under.

Tracked since
Jul 1, 2026
Newest source mention
Sep 26, 2026

Where the results come from

Verification: 184 scores · 43 independently verified · 96 aggregator-attributed · 23 vendor cross-reference · 22 vendor-reported. How these tiers are assigned

From 17 sources on 11 sites. Artificial Analysis supplies 98 of them; the 43 independently verified results come from 6 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai98
  • arcprize.org26
  • api.llm-stats.com18
  • thinkingmachines.ai13
  • deepmind.google7
  • livebench.ai7
  • epoch.ai5
  • deploymentsafety.openai.com4
  • huggingface.co3
  • datasets-server.huggingface.co2
  • simple-bench.com1

Also known as

How our sources name GPT-5.6 Luna at each reasoning setting.

SettingShort formAPI id
lowgpt-5.6 luna (low) gpt-5.6 luna 2026-07-30 (low)gpt-5-6-luna-low
mediumgpt-5.6 luna (medium) gpt-5.6 luna 2026-07-30 (medium)gpt-5-6-luna-medium
highgpt-5.6 luna (high) gpt-5.6 luna 2026-07-30 (high)gpt-5-6-luna-high
xhighgpt-5.6 luna (xhigh) gpt-5.6 luna 2026-07-30 (xhigh)gpt-5-6-luna-xhigh gpt-5.6-luna-xhigh gpt-5.6-luna-xhigh (codex-harness)
maxgpt-5.6 luna (max) gpt-5.6 luna 2026-07-30 (max)—
Also listed asgpt-5-6-lunagpt-5-6-luna-non-reasoninggpt-5.6 luna (non-reasoning)GPT-5.6 Luna 2026-07-30 (None)gpt-5.6-luna:batchopenai.gpt-5.6-lunagpt-5.6 luna (none)chatgpt luna 5.6luna 5.6