GPT-5.6 Luna vs GPT-5.6 Terra
Wins 0 of 8 areas
—
Wins 8 of 8 areas
Coding · Agents · Reasoning · Facts · Images and charts · Math · Long documents · Following instructions
GPT-5.6 Terra is the stronger all-rounder.GPT-5.6 Luna is cheaper.
Scores updated · 54 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- ReasoningHard problems that need careful thinking06GPT-5.6 Terra6 of 7 tests · 1 tie
- AgentsCarrying out multi-step tasks on its own05GPT-5.6 Terra5 of 6 tests · 1 tie
- FactsGetting facts right instead of making them up03GPT-5.6 Terra3 of 3 tests
- Images and chartsUnderstanding pictures, charts and video03GPT-5.6 Terra3 of 3 tests
- MathCompetition and research-level math03GPT-5.6 Terra3 of 3 tests
- CodingWriting and fixing software13GPT-5.6 Terra3 of 6 tests · 2 ties
- Long documentsFinding answers in very long texts01GPT-5.6 Terra1 of 2 tests · 1 tie
- Following instructionsDoing exactly what it is asked01GPT-5.6 Terra1 of 1 test
Following instructions rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
GPT-5.6 Luna costs 90% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where GPT-5.6 Terra pulls ahead
- Abstract visual puzzles that people can solveARC-AGI-2+24.3points ahead
- Complex command-line tasks across many fieldsTerminal-Bench 4.0+23.8points ahead
- Finds one of eight look-alike replies in a long chatMRCR v2 8 needle 128k (average)+18.7points ahead
Where GPT-5.6 Luna pulls ahead
- Programming problems, refreshed regularlyLiveBench · Coding+4.6points ahead
Every test, side by side
All 54 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGPT-5.6 Terra
- Terminal-Bench 2.1GPT-5.6 Terra by 7.180.988+7.1
- LiveBench · Agentic CodingGPT-5.6 Terra by 6.648.455+6.6
- LiveBench · CodingGPT-5.6 Luna by 4.682.978.3+4.6
- SciCodeGPT-5.6 Terra by 1.453.655+1.4
- SWE-bench Protie62.763.4tie
- LMArena · WebDevtie15201522tie
AgentsGPT-5.6 Terra
- Terminal-Bench 4.0GPT-5.6 Terra by 23.811.635.4+23.8
- AA IT-Bench SREGPT-5.6 Terra by 10.740.351+10.7
- τ-Bench V3 · BankingGPT-5.6 Terra by 9.131.140.2+9.1
- BrowseCompGPT-5.6 Terra by 4.283.387.5+4.2
- AA ApexAgentsGPT-5.6 Terra by 3.135.838.9+3.1
- GDPValtie48.247.7tie
ReasoningGPT-5.6 Terra
- ARC-AGI-2GPT-5.6 Terra by 24.359.683.9+24.3
- CritPtGPT-5.6 Terra by 9.420.630+9.4
- LiveBench · ReasoningGPT-5.6 Terra by 585.690.6+5
- Humanity's Last ExamGPT-5.6 Terra by 3.439.542.9+3.4
- SimpleBenchGPT-5.6 Terra by 2.146.848.9+2.1
- GPQA DiamondGPT-5.6 Terra by 1.491.192.5+1.4
- ARC-AGI-3tie0.20.8tie
FactsGPT-5.6 Terra
- AA-Omniscience · Non-hallucinationGPT-5.6 Terra by 4.77.412.1+4.7
- AA-Omniscience · AccuracyGPT-5.6 Terra by 4.142.746.8+4.1
- SimpleQA VerifiedGPT-5.6 Terra by 2.24143.2+2.2
Images and chartsGPT-5.6 Terra
- CharXiv (RQ)GPT-5.6 Terra by 3.282.785.9+3.2
- MMMU-ProGPT-5.6 Terra by 2.178.680.7+2.1
- LMArena · VisionGPT-5.6 Terra by 11 rating points12601271+11 rating
MathGPT-5.6 Terra
- FrontierMath Tier 4GPT-5.6 Terra by 9.858.568.3+9.8
- LiveBench · MathematicsGPT-5.6 Terra by 7.787.294.9+7.7
- FrontierMath Tiers 1-3 (v2)GPT-5.6 Terra by 3.982.186+3.9
Long documentsGPT-5.6 Terra
- MRCR v2 8 needle 128k (average)GPT-5.6 Terra by 18.774.893.5+18.7
- AA-LCRtie83.783tie
Following instructionsGPT-5.6 Terra
- LiveBench · Instruction FollowingGPT-5.6 Terra by 4.560.164.6+4.5
Other results23 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceGPT-5.6 Terra by 10.4-10.30.1+10.4
- livebench_languageGPT-5.6 Terra by 10.372.682.9+10.3
- FrontierMath (overall)GPT-5.6 Terra by 6.378.684.9+6.3
- ARC-AGI-1GPT-5.6 Terra by 5.890.796.5+5.8
- GDPval-AA v2 EloGPT-5.6 Luna by 56 rating points15841528+56 rating
- Artificial Analysis Coding IndexGPT-5.6 Terra by 5.371.476.7+5.3
- AA IntelligenceGPT-5.6 Terra by 4.837.342.1+4.8
- OSWorld 2.0GPT-5.6 Terra by 4.645.650.2+4.6
- Terminal-Bench 4.0GPT-5.6 Terra by 4.217.321.5+4.2
- Artificial AnalysisGPT-5.6 Terra by 45155+4
- DeepSWE 1.1GPT-5.6 Terra by 36770+3
- DeepSWEGPT-5.6 Terra by 2.467.269.6+2.4
- GDP (Surge AI)GPT-5.6 Terra by 222.724.7+2
- HealthBench ProfessionalGPT-5.6 Terra by 255.757.7+2
- HealthBench Professional length-adjustedGPT-5.6 Terra by 255.757.7+2
- livebench_data_analysisGPT-5.6 Terra by 1.37879.3+1.3
- HealthBenchGPT-5.6 Terra by 1.255.857+1.2
- HealthBench length-adjustedGPT-5.6 Terra by 1.255.857+1.2
- AA Agentic IndexGPT-5.6 Terra by 142.743.7+1
- HealthBench Hard length-adjustedtie3232.7tie
- Toolathlontie53.453.1tie
- Agents' Last Examtie50.350.4tie
- HealthBench Consensus length-adjustedtie95.195.1tie
Questions people ask
Which is better, GPT-5.6 Luna or GPT-5.6 Terra?
GPT-5.6 Terra wins all eight areas we test: coding, agents, reasoning, facts, images and charts, math, long documents and following instructions. GPT-5.6 Luna wins none, but costs 90% less.
Which is better for coding?
GPT-5.6 Terra. It wins 3 of the 6 coding tests both models report; GPT-5.6 Luna wins 1, and 2 are ties.
Which is cheaper?
GPT-5.6 Luna costs $0.20 per million input tokens and $1.20 per million output tokens; GPT-5.6 Terra costs $2.00 and $12.00. That makes GPT-5.6 Luna about 90% cheaper for the same work.
How do you compare the two?
We use the 54 benchmark tests both models have published scores on. The verdict counts the 31 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 23 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.