GPT-5.6 Luna
GPT-5.6 Luna is strong in long context; capable in reasoning, coding, agentic tasks, and multimodal tasks; and behind the leaders in math and factuality. Too few results yet to rate safety, multilingual tasks, or instruction following.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multilingual or Instruction Following.
Price
$0.20input$1.20outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
184results on65benchmarks
- 43 independently verified
- 96 aggregator
- 22 vendor-reported
- 23 cross-referenced
From 17 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
25 papers reference GPT-5.6 LunaGPT-5.6 Luna benchmark results
184 results on 65 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
5.0% behind the leader2 of 3 ranked benchmarks measured
- 83.67Oct 8, 2026
- MRCR v2 (8-needle, 128K)74.80Jul 22, 2026GDM-MRCR v2 (8-needle) (128k (average))
16.2% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond91.11Oct 8, 2026gpqa
- LiveBench · Reasoning85.64Oct 8, 2026livebench_reasoning@2026-06-25
- 20.57Oct 8, 2026
- 59.58Aug 7, 2026
- Humanity's Last Exam39.48Oct 8, 2026aa_hle
- 46.80Sep 21, 2026
Show 34 more reasoning resultsHide 34 reasoning results
- 0.83Sep 22, 2026
- 43.47Sep 22, 2026
- 3.61Sep 21, 2026
- 59.54Sep 21, 2026
- 47.64Sep 21, 2026
- 29.31Sep 21, 2026
- 7.36Sep 21, 2026
- 5.14Sep 21, 2026
- 7.36Aug 11, 2026
- 33.33Aug 9, 2026
- 47.60Jul 30, 2026
- 0.18Sep 21, 2026
- 0.02Sep 21, 2026
- 0.10Sep 21, 2026
- 0.17Sep 21, 2026
- 0.18Oct 7, 2026
- 0.29Oct 8, 2026
- 16.57Oct 8, 2026
- 4.86Oct 8, 2026
- 2.57Oct 8, 2026
- 20.60Jul 30, 2026
- GPQA Diamond64.55Oct 8, 2026gpqa
- GPQA Diamond89.19Oct 8, 2026gpqa
- GPQA Diamond89.49Oct 8, 2026gpqa
- GPQA Diamond85.86Oct 8, 2026gpqa
- GPQA Diamond83.54Oct 8, 2026gpqa
- GPQA Diamond92.30Oct 7, 2026GPQA
- 89.50Jul 30, 2026
- Humanity's Last Exam7.23Oct 8, 2026aa_hle
- Humanity's Last Exam36.98Oct 8, 2026aa_hle
- Humanity's Last Exam33.41Oct 8, 2026aa_hle
- Humanity's Last Exam25.76Oct 8, 2026aa_hle
- Humanity's Last Exam19.83Oct 8, 2026aa_hle
- Humanity's Last Exam35.60Jul 30, 2026HLE text only
20.7% behind the leader7 of 10 ranked benchmarks measured
- 93.00Jul 30, 2026
- LiveBench · Coding82.92Oct 8, 2026livebench_coding@2026-06-25
- Terminal-Bench 2.180.90Oct 8, 2026terminalbenchV21
- SciCode53.59Oct 8, 2026aa_scicode
- 1519.92Sep 21, 2026
- 62.70Oct 7, 2026
- LiveBench · Agentic Coding48.43Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 14 more coding resultsHide 14 coding results
- SciCode40.39Oct 8, 2026aa_scicode
- SciCode51.62Oct 8, 2026aa_scicode
- SciCode50.46Oct 8, 2026aa_scicode
- SciCode46.76Oct 8, 2026aa_scicode
- SciCode46.06Oct 8, 2026aa_scicode
- 50.00Jul 30, 2026
- SWE-bench Pro62.70Jul 22, 2026SWE-Bench Pro (Public)
- Terminal-Bench 2.138.95Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.169.66Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.177.90Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.153.18Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.143.45Oct 8, 2026terminalbenchV21
- 84.70Oct 7, 2026
- 82.50Jul 30, 2026
21.0% behind the leader5 of 7 ranked benchmarks measured
- 83.30Oct 7, 2026
- 72.60Jul 22, 2026
- τ-Bench V3 · Banking31.13Oct 8, 2026tauBanking
- 48.19Oct 8, 2026
- Terminal-Bench 4.011.62Oct 8, 2026
Show 17 more agentic resultsHide 17 agentic results
- 35.84Oct 8, 2026
- 40.32Oct 8, 2026
- 21.12Oct 8, 2026
- 41.66Oct 8, 2026
- 45.14Oct 8, 2026
- 31.56Oct 8, 2026
- 24.85Oct 8, 2026
- Terminal-Bench 4.01.01Oct 8, 2026
- Terminal-Bench 4.02.53Oct 8, 2026
- Terminal-Bench 4.03.54Oct 8, 2026
- Terminal-Bench 4.00.51Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking9.69Oct 8, 2026tauBanking
- τ-Bench V3 · Banking25.15Oct 8, 2026tauBanking
- τ-Bench V3 · Banking28.66Oct 8, 2026tauBanking
- τ-Bench V3 · Banking17.73Oct 8, 2026tauBanking
- τ-Bench V3 · Banking12.78Oct 8, 2026tauBanking
22.8% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro78.55Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)82.70Jul 22, 2026CharXiv Reasoning (No tools)
26.9% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics87.20Oct 8, 2026livebench_math@2026-06-25
- 82.11Sep 21, 2026
- FrontierMath Tier 458.50Oct 7, 2026FrontierMath Tier 4 (v2)
Show 5 more math resultsHide 5 math results
- 97.60Jul 30, 2026
- 60.98Sep 21, 2026
- 39.65Sep 22, 2026
- 41.40Sep 21, 2026
- 98.50Jul 30, 2026
36.7% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy42.73Oct 8, 2026omniscienceAccuracy
- 41.00Sep 21, 2026
- AA-Omniscience · Non-hallucination7.42Oct 8, 2026omniscienceNonHallucination
Show 10 more factuality resultsHide 10 factuality results
- AA-Omniscience · Accuracy28.60Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy42.45Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy41.78Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy40.70Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy39.62Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination25.00Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination7.62Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination7.53Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination9.13Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination10.13Oct 8, 2026omniscienceNonHallucination
0 of 3 ranked benchmarks measured
- LiveBench · Instruction Following60.12Oct 8, 2026livebench_instruction_following@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis78.03Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language72.57Oct 8, 2026livebench_language@2026-06-25
- 3.33Sep 22, 2026
- AA Intelligence37.00Sep 21, 2026Artificial Analysis Intelligence Index
- 39.50Sep 21, 2026
- 88.00Sep 21, 2026
Show 58 more resultsHide 58 results
- 15.17Sep 9, 2026
- 39.45Sep 9, 2026
- 35.62Sep 9, 2026
- 42.66Sep 9, 2026
- 25.42Sep 9, 2026
- 17.88Sep 9, 2026
- AA Intelligence52.00Aug 29, 2026Artificial Analysis Intelligence Index
- AA Intelligence15.53Oct 8, 2026aa_intelligence_index
- AA Intelligence34.56Oct 8, 2026aa_intelligence_index
- AA Intelligence32.12Oct 8, 2026aa_intelligence_index
- AA Intelligence25.04Oct 8, 2026aa_intelligence_index
- AA Intelligence37.32Oct 8, 2026aa_intelligence_index
- AA Intelligence21.01Oct 8, 2026aa_intelligence_index
- -24.95Oct 8, 2026
- -12.00Oct 8, 2026
- -10.77Oct 8, 2026
- -10.28Oct 8, 2026
- -13.18Oct 8, 2026
- -14.65Oct 8, 2026
- 50.30Oct 7, 2026
- 87.67Sep 21, 2026
- 76.50Sep 21, 2026
- 56.50Sep 21, 2026
- 34.17Sep 21, 2026
- 54.00Aug 11, 2026
- 78.00Aug 9, 2026
- 85.00Aug 8, 2026
- 90.67Aug 7, 2026
- 87.70Jul 30, 2026
- 51.00Oct 7, 2026
- Artificial Analysis Coding Index39.28Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index63.34Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index68.60Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index50.73Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index71.45Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index44.16Sep 9, 2026aa_coding_index
- BrowseComp (context management)84.00Jul 30, 2026BrowseComp with context management
- 67.20Oct 7, 2026
- 67.00Oct 7, 2026
- DeepSWE 1.167.00Jul 22, 2026DeepSWE v1.1
- FrontierMath (overall)78.60Oct 7, 2026FrontierMath
- GDP (Surge AI)22.70Oct 7, 2026GDP.pdf
- 1530.00Jul 30, 2026
- GDPval-AA v2 Elo1584.00Jul 22, 2026GDPVal-AA v2 (Elo)
- 55.80Oct 7, 2026
- 95.10Sep 29, 2026
- 32.00Sep 29, 2026
- 55.80Sep 29, 2026
- 55.70Oct 7, 2026
- 55.70Sep 29, 2026
- HLE (with tools)48.90Jul 30, 2026HLE with tools
- 47.60Jul 22, 2026
- 45.60Oct 7, 2026
- 62.70Jul 30, 2026
- Terminal-Bench 4.017.30Oct 7, 2026
- 53.40Oct 7, 2026
- 67.90Jul 30, 2026
- τ³-Bench Banking24.30Jul 30, 2026Tau 3 Banking
GPT-5.6 Luna: common questions
Who makes GPT-5.6 Luna?
GPT-5.6 Luna is made by OpenAI.
When was GPT-5.6 Luna released?
GPT-5.6 Luna was released on Jul 9, 2026, according to Artificial Analysis.
What is GPT-5.6 Luna good at?
GPT-5.6 Luna is strong in long context; capable in reasoning, coding, agentic tasks, and multimodal tasks; and behind the leaders in math and factuality. Too few results yet to rate safety, multilingual tasks, or instruction following.
How much does GPT-5.6 Luna cost?
GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it is cheaper than 60% of the 331 priced models we track.
How many benchmarks has GPT-5.6 Luna been tested on?
We track 184 results for GPT-5.6 Luna on 65 benchmarks from 17 sources, 43 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.6 Luna support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.6 Luna.
About this record
Where GPT-5.6 Luna's numbers come from, and every name it appears under.
- Tracked since
- Jul 1, 2026
- Newest source mention
- Sep 26, 2026
Where the results come from
Verification: 184 scores · 43 independently verified · 96 aggregator-attributed · 23 vendor cross-reference · 22 vendor-reported. How these tiers are assigned
From 17 sources on 11 sites. Artificial Analysis supplies 98 of them; the 43 independently verified results come from 6 sites. Bars are coloured by trust tier.
- artificialanalysis.ai98
- arcprize.org26
- api.llm-stats.com18
- thinkingmachines.ai13
- deepmind.google7
- livebench.ai7
- epoch.ai5
- deploymentsafety.openai.com4
- huggingface.co3
- datasets-server.huggingface.co2
- simple-bench.com1
Also known as
How our sources name GPT-5.6 Luna at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | gpt-5.6 luna (low) gpt-5.6 luna 2026-07-30 (low) | gpt-5-6-luna-low |
| medium | gpt-5.6 luna (medium) gpt-5.6 luna 2026-07-30 (medium) | gpt-5-6-luna-medium |
| high | gpt-5.6 luna (high) gpt-5.6 luna 2026-07-30 (high) | gpt-5-6-luna-high |
| xhigh | gpt-5.6 luna (xhigh) gpt-5.6 luna 2026-07-30 (xhigh) | gpt-5-6-luna-xhigh gpt-5.6-luna-xhigh gpt-5.6-luna-xhigh (codex-harness) |
| max | gpt-5.6 luna (max) gpt-5.6 luna 2026-07-30 (max) | — |