GPT-6 Luna
GPT-6 Luna is strong in long context; capable in reasoning and factuality; and behind the leaders in multimodal tasks, agentic tasks, coding, and math. Too few results yet to rate safety, multilingual tasks, or instruction following.
- Price per million tokens
$0.10input$0.50output
Price from Artificial Analysis · 3 providers tracked · All prices
Cheaper than 80% of 323 priced models · 3:1 input-to-output blend, log scale - Evidence
121results on44benchmarks
- 33 independently verified
- 66 aggregator
- 3 vendor-reported
- 19 cross-referenced
From 10 sources · latest Sep 29, 2026 · How verification works
- API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Capability profile
Bars show the model's median result as a share of the leading model's, per capability.
Too few results to rate: SafetyMultilingualInstruction Following
GPT-6 Luna benchmark results
121 results on 44 benchmarks, grouped by capability. Every score links to its source; a bar is the result as a share of the capability leader's.
Long Context
Strong1 of 3 ranked benchmarks measuredFull long context ranking- AA-LCR83.3394%aa_lcr
Reasoning
Capable4 of 6 ranked benchmarks measuredFull reasoning ranking- LiveBench · Reasoning81.7788%livebench_reasoning@2026-06-25
- 59.3162%
- 19.4360%
- Humanity's Last Exam38.5160%aa_hle
27 more reasoning results
- 0.00—
- 4.58—
- 18.06—
- 31.39—
- 41.94—
- 0.05—
- 0.24—
- 0.32—
- 0.38—
- 0.51—
- 0.59—
- 0.03—
- 0.03—
- 0.19—
- 0.18—
- 0.16—
- 0.10—
- 2.57—
- 1.14—
- 15.43—
- 10.57—
- 17.43—
- Humanity's Last Exam20.25—aa_hle
- Humanity's Last Exam8.57—aa_hle
- Humanity's Last Exam32.95—aa_hle
- Humanity's Last Exam28.27—aa_hle
- Humanity's Last Exam34.34—aa_hle
Factuality
Capable2 of 4 ranked benchmarks measuredFull factuality ranking- AA-Omniscience · Accuracy43.7865%omniscienceAccuracy
- AA-Omniscience · Non-hallucination23.2726%omniscienceNonHallucination
10 more factuality results
- AA-Omniscience · Accuracy40.77—omniscienceAccuracy
- AA-Omniscience · Accuracy42.80—omniscienceAccuracy
- AA-Omniscience · Accuracy31.85—omniscienceAccuracy
- AA-Omniscience · Accuracy44.17—omniscienceAccuracy
- AA-Omniscience · Accuracy43.13—omniscienceAccuracy
- AA-Omniscience · Non-hallucination15.70—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination21.13—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination15.56—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination15.27—omniscienceNonHallucination
- AA-Omniscience · Non-hallucination17.61—omniscienceNonHallucination
Multimodal
Limited1 of 6 ranked benchmarks measuredFull multimodal ranking- MMMU-Pro75.5586%aa_mmmu_pro
Agentic
Limited2 of 7 ranked benchmarks measuredFull agentic ranking- 43.3549%
- Terminal-Bench 4.012.6320%
10 more agentic results
- 24.61—
- 39.52—
- 26.70—
- 39.84—
- 35.92—
- Terminal-Bench 4.00.00—
- Terminal-Bench 4.04.55—
- Terminal-Bench 4.01.52—
- Terminal-Bench 4.02.53—
- Terminal-Bench 4.08.08—
Coding
Limited4 of 10 ranked benchmarks measuredFull coding ranking- LiveBench · Coding78.9586%livebench_coding@2026-06-25
- 1592.6585%
- SciCode54.6382%aa_scicode
- LiveBench · Agentic Coding51.2166%livebench_agentic_coding@2026-06-25
Math
Limited1 of 5 ranked benchmarks measuredFull math ranking- LiveBench · Mathematics89.1292%livebench_math@2026-06-25
Instruction Following
Not enough data0 of 3 ranked benchmarks measuredFull instruction following ranking- LiveBench · Instruction Following55.93—livebench_instruction_following@2026-06-25
More results
- livebench_language73.83—livebench_language@2026-06-25
- livebench_data_analysis73.37—livebench_data_analysis@2026-06-25
- 8.83—
- 37.67—
- 61.00—
- 70.33—
37 more results
- AA Intelligence37.00—Artificial Analysis Intelligence Index
- AA Intelligence20.92—aa_intelligence_index
- AA Intelligence18.26—aa_intelligence_index
- AA Intelligence37.26—aa_intelligence_index
- AA Intelligence32.15—aa_intelligence_index
- AA Intelligence29.46—aa_intelligence_index
- AA Intelligence33.88—aa_intelligence_index
- AA-Omniscience-9.17—aa_omniscience
- AA-Omniscience-5.50—aa_omniscience
- AA-Omniscience0.65—aa_omniscience
- AA-Omniscience-21.90—aa_omniscience
- AA-Omniscience-5.05—aa_omniscience
- AA-Omniscience-1.83—aa_omniscience
- Agentic Safe Completions - Chat prod - Chat Plugins0.85—
- 50.90—
- 73.00—
- 86.67—
- 66.60—
- Dynamic Benchmarks - Emotional reliance0.96—
- Dynamic Benchmarks - Mental health1.00—
- Dynamic Benchmarks - Self-harm0.92—
- 95.90—
- 31.40—
- HealthBench length-adjusted54.50—HealthBench length-adjusted
- 60.80—
- Image input evaluations - extremism0.98—
- Image input evaluations - harms-erotic0.99—
- Image input evaluations - hate1.00—
- Image input evaluations - self-harm1.00—
- 52.70—
- Production Benchmarks - Gore0.88—
- U18 evaluations - Age-restricted goods, services, and dangerous challenges / activities0.85—
- U18 evaluations - Eating Disorders0.87—
- U18 evaluations - Emotional Reliance0.95—
- U18 evaluations - Gore0.88—
- U18 evaluations - Self Harm0.98—
- U18 evaluations - Sexual Content0.95—
About this record
Verification: 121 scores · 33 independently verified · 66 aggregator-attributed · 19 vendor cross-reference · 3 vendor-reported. How these tiers are assigned
- Also listed as
- gpt-6-luna-xhighgpt-6-luna-mediumgpt-6-luna-non-reasoninggpt-6-luna-highgpt-6-luna-lowgpt-6 luna (xhigh)gpt-6 luna (medium)gpt-6 luna (non-reasoning)gpt-6 luna (max)gpt-6 luna (high)gpt-6 luna (low)gpt-6 luna (none)gpt-6 luna - provider adapter (max)gpt-6 luna - provider adapter (xhigh)gpt-6 luna - provider adapter (high)gpt-6-luna-maxand 2 more
- Tracked since
- Sep 22, 2026
- Newest source mention
- Sep 27, 2026