GPT-5.6 Terra vs Claude Opus 5
Wins 1 of 8 areas
Long documents
Wins 6 of 8 areas
Coding · Agents · Reasoning · Facts · Images and charts · Math
Claude Opus 5 is the stronger all-rounder.GPT-5.6 Terra is cheaper and better at long documents.
Both rank among the ten best models we track in coding.
Scores updated · 48 tests both models report · How we compare
Where each one wins
Tests won in each of the eight areas we test. Each piece is one test, so a longer bar means more evidence; grey means the two scored within a point of each other.
- CodingWriting and fixing software06Claude Opus 56 of 6 tests
- ReasoningHard problems that need careful thinking04Claude Opus 54 of 7 tests · 3 ties
- AgentsCarrying out multi-step tasks on its own04Claude Opus 54 of 4 tests
- FactsGetting facts right instead of making them up03Claude Opus 53 of 3 tests
- Images and chartsUnderstanding pictures, charts and video12Claude Opus 52 of 3 tests
- MathCompetition and research-level math01Claude Opus 51 of 3 tests · 2 ties
- Long documentsFinding answers in very long texts10GPT-5.6 Terra1 of 1 test
- Following instructionsDoing exactly what it is asked00Even0 each · 1 tie
Long documents and following instructions rest on a single test each.
What it costs
Prices per million tokens, roughly 750,000 words. The bars show the cost of a million tokens read plus a million written.
GPT-5.6 Terra costs 53% less for the same work.
The biggest differences
The tests each model wins by the widest margin, up to three each. Scores are out of 100.
Where Claude Opus 5 pulls ahead
- Common-sense trick questionsSimpleBench+31.7points ahead
- Solves new puzzle games it has never seenARC-AGI-3+29.4points ahead
- Avoids making up answers it doesn't knowAA-Omniscience · Non-hallucination+27.1points ahead
Where GPT-5.6 Terra pulls ahead
- Reasons across sets of long documentsAA-LCR+3.7points ahead
- Reasoning about charts from research papersCharXiv (RQ)+2.2points ahead
Every test, side by side
All 48 tests both models report. The winning score is in its model's colour; marks a score checked independently.
CodingClaude Opus 5
- LMArena · WebDevClaude Opus 5 by 170 rating points15221692+170 rating
- SWE-bench ProClaude Opus 5 by 15.863.479.2+15.8
- LiveBench · Agentic CodingClaude Opus 5 by 10.25565.2+10.2
- LiveBench · CodingClaude Opus 5 by 3.278.381.5+3.2
- SciCodeClaude Opus 5 by 1.45556.4+1.4
- Terminal-Bench 2.1Claude Opus 5 by 1.18889.1+1.1
AgentsClaude Opus 5
- Terminal-Bench 4.0Claude Opus 5 by 13.635.449+13.6
- GDPValClaude Opus 5 by 13.547.761.2+13.5
- BrowseCompClaude Opus 5 by 3.387.590.8+3.3
- τ-Bench V3 · BankingClaude Opus 5 by 1.940.242.1+1.9
ReasoningClaude Opus 5
- SimpleBenchClaude Opus 5 by 31.748.980.6+31.7
- ARC-AGI-3Claude Opus 5 by 29.40.830.2+29.4
- Humanity's Last ExamClaude Opus 5 by 1242.954.9+12
- ARC-AGI-2Claude Opus 5 by 6.583.990.4+6.5
- CritPttie3029.1tie
- GPQA Diamondtie92.593.2tie
- LiveBench · Reasoningtie90.691.2tie
FactsClaude Opus 5
- AA-Omniscience · Non-hallucinationClaude Opus 5 by 27.112.139.2+27.1
- SimpleQA VerifiedClaude Opus 5 by 16.743.259.9+16.7
- AA-Omniscience · AccuracyClaude Opus 5 by 14.146.860.9+14.1
Images and chartsClaude Opus 5
- LMArena · VisionClaude Opus 5 by 50 rating points12711321+50 rating
- MMMU-ProClaude Opus 5 by 480.784.7+4
- CharXiv (RQ)GPT-5.6 Terra by 2.285.983.7+2.2
MathClaude Opus 5
- FrontierMath Tier 4Claude Opus 5 by 4.968.373.2+4.9
- LiveBench · Mathematicstie94.995.7tie
- FrontierMath Tiers 1-3 (v2)tie8685.6tie
Following instructionsEven
- LiveBench · Instruction Followingtie64.663.8tie
Other results20 tests, not counted
Tests outside the eight areas. They are not counted above: several are summary scores built from other tests, or the same test under another name.
- GDP (Surge AI)Claude Opus 5 by 60.824.785.5+60.8
- AA-OmniscienceClaude Opus 5 by 370.137.1+37
- Terminal-Bench 4.0Claude Opus 5 by 30.321.551.8+30.3
- GDPval-AA v2 EloClaude Opus 5 by 296 rating points15281824+296 rating
- OSWorld 2.0 (partial)Claude Opus 5 by 25.250.275.4+25.2
- OSWorld 2.0Claude Opus 5 by 20.450.270.6+20.4
- Agents' Last ExamGPT-5.6 Terra by 18.850.431.6+18.8
- AA Agentic IndexClaude Opus 5 by 12.543.756.2+12.5
- AA IntelligenceClaude Opus 5 by 8.742.150.8+8.7
- GDP.PDF (All pass rate)Claude Opus 5 by 82937+8
- livebench_languageClaude Opus 5 by 5.882.988.7+5.8
- livebench_data_analysisGPT-5.6 Terra by 4.779.374.5+4.7
- LVBenchGPT-5.6 Terra by 3.578.975.4+3.5
- HLE-VerifiedClaude Opus 5 by 3.351.154.4+3.3
- HealthBench ProfessionalClaude Opus 5 by 2.157.759.8+2.1
- HealthBench Professional length-adjustedClaude Opus 5 by 2.157.759.8+2.1
- Artificial Analysis Coding IndexClaude Opus 5 by 1.376.778+1.3
- DeepSWE 1.1GPT-5.6 Terra by 1.27068.8+1.2
- ARC-AGI-1Claude Opus 5 by 196.597.5+1
- HealthBench length-adjustedtie5757.8tie
Questions people ask
Which is better, GPT-5.6 Terra or Claude Opus 5?
Claude Opus 5 wins six of the eight areas we test: coding, agents, reasoning, facts, images and charts and math. GPT-5.6 Terra wins long documents, and costs 53% less. They are level on following instructions.
Which is better for coding?
Claude Opus 5. It wins 6 of the 6 coding tests both models report; GPT-5.6 Terra wins none.
Which is cheaper?
GPT-5.6 Terra costs $2.00 per million input tokens and $12.00 per million output tokens; Claude Opus 5 costs $5.00 and $25.00. That makes GPT-5.6 Terra about 53% cheaper for the same work.
How do you compare the two?
We use the 48 benchmark tests both models have published scores on. The verdict counts the 28 tests in the eight capability areas, and a gap under one point (ten on rating-style scales) counts as a tie. The other 20 are listed but not counted, because several are summary scores or repeat a test. Each score is the one shown on the model's own page.