Gemini 3.5 Flash Lite vs GPT-5.6 Luna
Wins 1 of 8 areas
Following instructions
Wins 6 of 8 areas
Coding · Agents · Reasoning · Images and charts · Math · Long documents
GPT-5.6 Luna is the stronger all-rounder.Gemini 3.5 Flash Lite is better at following instructions.
Scores updated · 41 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software07GPT-5.6 Luna7 of 7 tests
- ReasoningHard problems that need careful thinking05GPT-5.6 Luna5 of 5 tests
- MathCompetition and research-level math05GPT-5.6 Luna5 of 5 tests
- AgentsCarrying out multi-step tasks on its own13GPT-5.6 Luna3 of 4 tests
- Images and chartsUnderstanding pictures, charts and video01GPT-5.6 Luna1 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts01GPT-5.6 Luna1 of 1 test
- Following instructionsDoing exactly what it is asked10Gemini 3.5 Flash Lite1 of 1 test
- FactsGetting facts right instead of making them up11Even1 each
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
GPT-5.6 Luna costs 50% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where GPT-5.6 Luna pulls ahead
- Unpublished advanced math problemsFrontierMath Tiers 1-3 (v2)+56.1points ahead
- Abstract visual puzzles that people can solveARC-AGI-2+49.3points ahead
- Harvard-MIT high-school math contest problemsHMMT Feb. 2026+34.9points ahead
Where Gemini 3.5 Flash Lite pulls ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+58.2points ahead
- Follows detailed instructions exactlyLiveBench · Instruction Following+7.1points ahead
- Completes tasks by operating a computer desktopOSWorld-Verified+1.4points ahead
Every test, side by side
All 41 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGPT-5.6 Luna
- Terminal-Bench 2.1GPT-5.6 Luna by 27.353.680.9+27.3
- SWE-bench VerifiedGPT-5.6 Luna by 187593+18
- SciCodeGPT-5.6 Luna by 12.341.353.6+12.3
- SWE-bench ProGPT-5.6 Luna by 8.554.262.7+8.5
- LMArena · WebDevGPT-5.6 Luna by 80 rating points14401520+80 rating
- LiveBench · CodingGPT-5.6 Luna by 6.876.182.9+6.8
- LiveBench · Agentic CodingGPT-5.6 Luna by 3.145.348.4+3.1
AgentsGPT-5.6 Luna
- GDPValGPT-5.6 Luna by 23.924.348.2+23.9
- τ-Bench V3 · BankingGPT-5.6 Luna by 13.617.531.1+13.6
- Terminal-Bench 4.0GPT-5.6 Luna by 10.6111.6+10.6
- OSWorld-VerifiedGemini 3.5 Flash Lite by 1.47472.6+1.4
ReasoningGPT-5.6 Luna
- ARC-AGI-2GPT-5.6 Luna by 49.310.359.6+49.3
- LiveBench · ReasoningGPT-5.6 Luna by 25.460.285.6+25.4
- Humanity's Last ExamGPT-5.6 Luna by 20.718.839.5+20.7
- CritPtGPT-5.6 Luna by 20.6020.6+20.6
- GPQA DiamondGPT-5.6 Luna by 7.383.891.1+7.3
FactsEven
- AA-Omniscience · Non-hallucinationGemini 3.5 Flash Lite by 58.265.67.4+58.2
- AA-Omniscience · AccuracyGPT-5.6 Luna by 13.229.542.7+13.2
Images and chartsGPT-5.6 Luna
- CharXiv (RQ)GPT-5.6 Luna by 6.276.582.7+6.2
- LMArena · Visiontie12661260tie
- MMMU-Protie7978.6tie
MathGPT-5.6 Luna
- FrontierMath Tier 4GPT-5.6 Luna by 58.5058.5+58.5
- FrontierMath Tiers 1-3 (v2)GPT-5.6 Luna by 56.12682.1+56.1
- HMMT Feb. 2026GPT-5.6 Luna by 34.963.698.5+34.9
- AIME 2026GPT-5.6 Luna by 15.482.297.6+15.4
- LiveBench · MathematicsGPT-5.6 Luna by 13.573.787.2+13.5
Following instructionsGemini 3.5 Flash Lite
- LiveBench · Instruction FollowingGemini 3.5 Flash Lite by 7.167.260.1+7.1
Other results13 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDPval-AA v2GPT-5.6 Luna by 390 rating points11401530+390 rating
- ARC-AGI-1GPT-5.6 Luna by 37.253.590.7+37.2
- AA Agentic IndexGPT-5.6 Luna by 26.815.942.7+26.8
- livebench_data_analysisGPT-5.6 Luna by 24.753.378+24.7
- Artificial Analysis Coding IndexGPT-5.6 Luna by 22.149.371.4+22.1
- AA-OmniscienceGemini 3.5 Flash Lite by 15.55.2-10.3+15.5
- AA IntelligenceGPT-5.6 Luna by 15.122.237.3+15.1
- Toolathlon VerifiedGPT-5.6 Luna by 10.857.167.9+10.8
- SWEBench Pro PublicGPT-5.6 Luna by 8.554.262.7+8.5
- MLE-BenchGPT-5.6 Luna by 8.439.247.6+8.4
- τ³-Bench BankingGPT-5.6 Luna by 7.816.524.3+7.8
- HLE (with tools)GPT-5.6 Luna by 6.442.548.9+6.4
- livebench_languagetie71.872.6tie
Questions people ask
Which is better, Gemini 3.5 Flash Lite or GPT-5.6 Luna?
GPT-5.6 Luna wins six of the eight areas we test: coding, agents, reasoning, images and charts, math and long documents. Gemini 3.5 Flash Lite wins following instructions. They are level on facts.
Which is better for coding?
GPT-5.6 Luna. It wins 7 of the 7 coding tests both models report; Gemini 3.5 Flash Lite wins none.
Which is cheaper?
Gemini 3.5 Flash Lite costs $0.30 per million input tokens and $2.50 per million output tokens; GPT-5.6 Luna costs $0.20 and $1.20. That makes GPT-5.6 Luna about 50% cheaper for the same work.
How do you compare the two?
We use the 41 benchmark tests both models have published scores on. The verdict counts the 28 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 13 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.