Gemini 3.7 Flash vs GPT-5.6 Terra
Wins 4 of 8 areas
Coding · Facts · Images and charts · Following instructions
Wins 2 of 8 areas
Agents · Math
Gemini 3.7 Flash wins more areas, narrowly.GPT-5.6 Terra is better at agents.
Both rank among the ten best models we track in long documents.
Scores updated · 46 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- FactsGetting facts right instead of making them up30Gemini 3.7 Flash3 of 3 tests
- Images and chartsUnderstanding pictures, charts and video30Gemini 3.7 Flash3 of 3 tests
- CodingWriting and fixing software31Gemini 3.7 Flash3 of 5 tests · 1 tie
- Following instructionsDoing exactly what it is asked10Gemini 3.7 Flash1 of 1 test
- AgentsCarrying out multi-step tasks on its own03GPT-5.6 Terra3 of 3 tests
- MathCompetition and research-level math03GPT-5.6 Terra3 of 3 tests
- ReasoningHard problems that need careful thinking22Even2 each · 1 tie
- Long documentsFinding answers in very long texts11Even1 each
Following instructions rests on a single test.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
Gemini 3.7 Flash costs 68% less for the same work.
The biggest differences
The three tests each model wins by the widest margin. Scores are out of 100.
Where Gemini 3.7 Flash pulls ahead
- Short factual questions, answered correctlySimpleQA Verified+26points ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+23.4points ahead
- Follows detailed instructions exactlyLiveBench · Instruction Following+15.3points ahead
Where GPT-5.6 Terra pulls ahead
- Extremely hard research-level math problemsFrontierMath Tier 4+31.7points ahead
- Complex command-line tasks across many fieldsTerminal-Bench 4.0+21.8points ahead
- Research-level physics problemsCritPt+15.7points ahead
Every test, side by side
All 46 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingGemini 3.7 Flash
- LMArena · WebDevGemini 3.7 Flash by 73 rating points15921519+73 rating
- LiveBench · Agentic CodingGemini 3.7 Flash by 3.358.355+3.3
- SciCodeGemini 3.7 Flash by 2.257.255+2.2
- Terminal-Bench 2.1GPT-5.6 Terra by 2.285.888+2.2
- LiveBench · Codingtie78.978.3tie
AgentsGPT-5.6 Terra
- Terminal-Bench 4.0GPT-5.6 Terra by 21.813.635.4+21.8
- τ-Bench V3 · BankingGPT-5.6 Terra by 7.432.840.2+7.4
- GDPValGPT-5.6 Terra by 3.144.647.7+3.1
ReasoningEven
- CritPtGPT-5.6 Terra by 15.714.330+15.7
- Humanity's Last ExamGemini 3.7 Flash by 547.942.9+5
- LiveBench · ReasoningGPT-5.6 Terra by 2.887.890.6+2.8
- GPQA DiamondGemini 3.7 Flash by 294.592.5+2
- ARC-AGI-2tie84.683.9tie
FactsGemini 3.7 Flash
- SimpleQA VerifiedGemini 3.7 Flash by 2669.243.2+26
- AA-Omniscience · Non-hallucinationGemini 3.7 Flash by 23.435.512.1+23.4
- AA-Omniscience · AccuracyGemini 3.7 Flash by 8.555.346.8+8.5
Images and chartsGemini 3.7 Flash
- MMMU-ProGemini 3.7 Flash by 4.885.580.7+4.8
- LMArena · VisionGemini 3.7 Flash by 44 rating points13151271+44 rating
- CharXiv (RQ)Gemini 3.7 Flash by 2.888.785.9+2.8
MathGPT-5.6 Terra
- FrontierMath Tier 4GPT-5.6 Terra by 31.736.668.3+31.7
- FrontierMath Tiers 1-3 (v2)GPT-5.6 Terra by 14.471.686+14.4
- LiveBench · MathematicsGPT-5.6 Terra by 1.493.594.9+1.4
Long documentsEven
- MRCR v2 8 needle 128k (average)Gemini 3.7 Flash by 3.59793.5+3.5
- AA-LCRGPT-5.6 Terra by 1.381.783+1.3
Following instructionsGemini 3.7 Flash
- LiveBench · Instruction FollowingGemini 3.7 Flash by 15.379.964.6+15.3
Other results21 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- AA-OmniscienceGemini 3.7 Flash by 26.426.50.1+26.4
- Agents' Last ExamGPT-5.6 Terra by 24.126.350.4+24.1
- livebench_data_analysisGPT-5.6 Terra by 11.36879.3+11.3
- Terminal-Bench 4.0GPT-5.6 Terra by 10.311.221.5+10.3
- GDP (Surge AI)Gemini 3.7 Flash by 9.33424.7+9.3
- AA Agentic IndexGPT-5.6 Terra by 7.336.443.7+7.3
- Code Arena (Elo)Gemini 3.7 Flash by 65 rating points15881523+65 rating
- LVBenchGemini 3.7 Flash by 6.585.478.9+6.5
- Terminal-bench 3.0GPT-5.6 Terra by 5.914.920.8+5.9
- GDP.PDF (All pass rate)Gemini 3.7 Flash by 53429+5
- DeepSWE 1.1GPT-5.6 Terra by 4.765.370+4.7
- GDPval-AA v2 EloGPT-5.6 Terra by 46 rating points14821528+46 rating
- AA IntelligenceGPT-5.6 Terra by 339.142.1+3
- livebench_languageGemini 3.7 Flash by 2.685.582.9+2.6
- HLE-VerifiedGemini 3.7 Flash by 2.553.651.1+2.5
- OSWorld 2.0GPT-5.6 Terra by 2.347.950.2+2.3
- Agent's Last Exam (Pass rate)GPT-5.6 Terra by 1.726.328+1.7
- ARC-AGI-1GPT-5.6 Terra by 195.596.5+1
- Artificial AnalysisGemini 3.7 Flash by 15655+1
- Artificial Analysis Coding Indextie76.176.7tie
- OSWorld 2.0 (partial)tie50.650.2tie
Questions people ask
Which is better, Gemini 3.7 Flash or GPT-5.6 Terra?
Gemini 3.7 Flash wins four of the eight areas we test: coding, facts, images and charts and following instructions. GPT-5.6 Terra wins agents and math. They are level on reasoning and long documents.
Which is better for coding?
Gemini 3.7 Flash. It wins 3 of the 5 coding tests both models report; GPT-5.6 Terra wins 1, and 1 is a tie.
Which is cheaper?
Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens; GPT-5.6 Terra costs $2.00 and $12.00. That makes Gemini 3.7 Flash about 68% cheaper for the same work.
How do you compare the two?
We use the 46 benchmark tests both models have published scores on. The verdict counts the 25 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 21 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.