GPT-5.6 Terra
GPT-5.6 Terra is at the frontier in long context; strong in reasoning; capable in coding, agentic tasks, multimodal tasks, and math; and behind the leaders in factuality and instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$2.00input$12.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
176results on65benchmarks
- 30 independently verified
- 110 aggregator
- 22 vendor-reported
- 14 cross-referenced
From 16 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
8 papers reference GPT-5.6 TerraGPT-5.6 Terra benchmark results
176 results on 65 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
Leads the field2 of 3 ranked benchmarks measured
- MRCR v2 (8-needle, 128K)93.50Aug 14, 2026GDM-MRCR v2 (8-needle) (128k (average))
- 83.00Oct 8, 2026
8.8% behind the leader6 of 6 ranked benchmarks measured
- LiveBench · Reasoning90.63Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond92.53Oct 8, 2026gpqa
- 30.00Oct 8, 2026
- 83.90Sep 21, 2026
- Humanity's Last Exam42.91Oct 8, 2026aa_hle
- 48.90Sep 21, 2026
Show 26 more reasoning resultsHide 26 reasoning results
- 74.17Sep 21, 2026
- 67.08Sep 21, 2026
- 37.50Sep 21, 2026
- 18.75Sep 21, 2026
- 0.80Sep 21, 2026
- 0.65Sep 21, 2026
- 0.49Sep 21, 2026
- 0.08Sep 21, 2026
- 0.01Sep 21, 2026
- 0.80Oct 7, 2026
- 9.43Oct 8, 2026
- 22.86Oct 8, 2026
- 27.14Oct 8, 2026
- 2.00Oct 8, 2026
- 17.43Oct 8, 2026
- GPQA Diamond84.34Oct 8, 2026gpqa
- GPQA Diamond74.65Oct 8, 2026gpqa
- GPQA Diamond89.60Oct 8, 2026gpqa
- GPQA Diamond90.81Oct 8, 2026gpqa
- GPQA Diamond87.17Oct 8, 2026gpqa
- GPQA Diamond92.90Oct 7, 2026GPQA
- Humanity's Last Exam29.19Oct 8, 2026aa_hle
- Humanity's Last Exam41.89Oct 8, 2026aa_hle
- Humanity's Last Exam38.51Oct 8, 2026aa_hle
- Humanity's Last Exam11.40Oct 8, 2026aa_hle
- Humanity's Last Exam33.27Oct 8, 2026aa_hle
12.1% behind the leader7 of 10 ranked benchmarks measured
- Terminal-Bench 2.188.01Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard57.58Oct 8, 2026aa_terminalbench_hard
- LiveBench · Coding78.25Oct 8, 2026livebench_coding@2026-06-25
- SciCode54.98Oct 8, 2026aa_scicode
- 1521.52Sep 21, 2026
- LiveBench · Agentic Coding54.95Oct 8, 2026livebench_agentic_coding@2026-06-25
- 63.40Oct 7, 2026
Show 14 more coding resultsHide 14 coding results
- SciCode49.88Oct 8, 2026aa_scicode
- SciCode52.43Oct 8, 2026aa_scicode
- SciCode45.14Oct 8, 2026aa_scicode
- SciCode52.31Oct 8, 2026aa_scicode
- SciCode50.46Oct 8, 2026aa_scicode
- Terminal-Bench 2.162.55Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.175.66Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.180.15Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.156.18Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.172.28Oct 8, 2026terminalbenchV21
- 87.40Oct 7, 2026
- 87.40Sep 11, 2026
- Terminal-Bench Hard43.94Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard62.88Oct 8, 2026aa_terminalbench_hard
15.6% behind the leader4 of 7 ranked benchmarks measured
- 87.50Oct 7, 2026
- τ-Bench V3 · Banking40.21Oct 8, 2026tauBanking
- Terminal-Bench 4.035.35Oct 8, 2026
- 47.66Oct 8, 2026
Show 16 more agentic resultsHide 16 agentic results
- 38.94Oct 8, 2026
- 51.04Oct 8, 2026
- 30.51Oct 8, 2026
- 29.98Oct 8, 2026
- 43.93Oct 8, 2026
- 47.46Oct 8, 2026
- 38.46Oct 8, 2026
- Terminal-Bench 4.01.52Oct 8, 2026
- Terminal-Bench 4.00.51Oct 8, 2026
- Terminal-Bench 4.010.10Oct 8, 2026
- Terminal-Bench 4.01.01Oct 8, 2026
- τ-Bench V3 · Banking18.76Oct 8, 2026tauBanking
- τ-Bench V3 · Banking15.67Oct 8, 2026tauBanking
- τ-Bench V3 · Banking29.69Oct 8, 2026tauBanking
- τ-Bench V3 · Banking28.66Oct 8, 2026tauBanking
- τ-Bench V3 · Banking25.57Oct 8, 2026tauBanking
18.9% behind the leader2 of 6 ranked benchmarks measured
- CharXiv (reasoning)85.90Aug 14, 2026CharXiv Reasoning (No tools)
- MMMU-Pro80.69Oct 8, 2026aa_mmmu_pro
Show 7 more multimodal resultsHide 7 multimodal results
20.7% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics94.91Oct 8, 2026livebench_math@2026-06-25
- 85.96Sep 21, 2026
- FrontierMath Tier 468.30Oct 7, 2026FrontierMath Tier 4 (v2)
Show 1 more math resultHide 1 math result
- 70.73Sep 21, 2026
30.4% behind the leader3 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy46.80Oct 8, 2026omniscienceAccuracy
- 43.20Sep 21, 2026
- AA-Omniscience · Non-hallucination12.12Oct 8, 2026omniscienceNonHallucination
Show 10 more factuality resultsHide 10 factuality results
- AA-Omniscience · Accuracy43.73Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy45.52Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy45.48Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy36.82Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy44.60Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination10.13Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination10.21Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination4.99Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination10.98Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination10.23Oct 8, 2026omniscienceNonHallucination
31.6% behind the leader2 of 3 ranked benchmarks measured
- IFBench71.22Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following64.62Oct 8, 2026livebench_instruction_following@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis79.31Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language82.89Oct 8, 2026livebench_language@2026-06-25
- 96.50Sep 21, 2026
- 94.00Sep 21, 2026
- 92.00Sep 21, 2026
- 77.00Sep 21, 2026
Show 58 more resultsHide 58 results
- 27.30Sep 9, 2026
- 43.73Sep 9, 2026
- 24.92Sep 9, 2026
- 42.19Sep 9, 2026
- 37.57Sep 9, 2026
- 31.44Sep 9, 2026
- AA Intelligence57.00Jul 10, 2026Artificial Analysis Intelligence Index
- AA Intelligence27.50Oct 8, 2026aa_intelligence_index
- AA Intelligence20.78Oct 8, 2026aa_intelligence_index
- AA Intelligence42.08Oct 8, 2026aa_intelligence_index
- AA Intelligence37.95Oct 8, 2026aa_intelligence_index
- AA Intelligence34.24Oct 8, 2026aa_intelligence_index
- AA Intelligence30.09Oct 8, 2026aa_intelligence_index
- -6.83Oct 8, 2026
- -2.98Oct 8, 2026
- 0.05Oct 8, 2026
- -3.47Oct 8, 2026
- -23.22Oct 8, 2026
- -5.13Oct 8, 2026
- 28.00Aug 14, 2026
- 50.40Oct 7, 2026
- 60.17Sep 21, 2026
- 1466Sep 13, 2026
- 55.00Oct 7, 2026
- Artificial Analysis Coding Index58.09Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index76.66Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index52.31Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index70.64Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index67.14Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index64.74Sep 9, 2026aa_coding_index
- 1523.00Aug 14, 2026
- 69.60Oct 7, 2026
- 70.00Oct 7, 2026
- DeepSWE 1.169.60Aug 14, 2026DeepSWE v1.1
- FrontierMath (overall)84.90Oct 7, 2026FrontierMath
- GDP (Surge AI)24.70Oct 7, 2026GDP.pdf
- GDP (Surge AI)24.70Aug 14, 2026GDP.pdf
- GDP.PDF (All pass rate)29.00Sep 9, 2026
- GDPval-AA v2 Elo1528.00Aug 14, 2026GDPVal-AA v2 (Elo)
- 57.00Oct 7, 2026
- 95.10Sep 29, 2026
- 32.70Sep 29, 2026
- 57.00Sep 29, 2026
- 57.70Oct 7, 2026
- 57.70Sep 29, 2026
- 51.10Aug 14, 2026
- 78.90Aug 14, 2026
- 50.20Oct 7, 2026
- OSWorld 2.0 (partial)50.20Sep 9, 2026OSWorld-2.0 (Partial score)
- 20.80Aug 14, 2026
- Terminal-Bench 4.021.50Oct 7, 2026
- Terminal-Bench 4.023.60Sep 11, 2026
- 53.10Oct 7, 2026
- τ²-Bench Telecom (AA run)60.53Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)78.36Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)86.26Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)80.41Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)72.81Oct 8, 2026aa_tau2
GPT-5.6 Terra: common questions
Who makes GPT-5.6 Terra?
GPT-5.6 Terra is made by OpenAI.
When was GPT-5.6 Terra released?
GPT-5.6 Terra was released on Jul 9, 2026, according to Artificial Analysis.
What is GPT-5.6 Terra good at?
GPT-5.6 Terra is at the frontier in long context; strong in reasoning; capable in coding, agentic tasks, multimodal tasks, and math; and behind the leaders in factuality and instruction following. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5.6 Terra cost?
GPT-5.6 Terra costs $2.00 per million input tokens and $12.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 85% of the 330 priced models we track.
How many benchmarks has GPT-5.6 Terra been tested on?
We track 176 results for GPT-5.6 Terra on 65 benchmarks from 16 sources, 30 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.6 Terra support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.6 Terra.
About this record
Where GPT-5.6 Terra's numbers come from, and every name it appears under.
- Tracked since
- Jul 1, 2026
- Newest source mention
- Sep 9, 2026
Where the results come from
Verification: 176 scores · 30 independently verified · 110 aggregator-attributed · 14 vendor cross-reference · 22 vendor-reported. How these tiers are assigned
From 16 sources on 10 sites. Artificial Analysis supplies 111 of them; the 30 independently verified results come from 7 sites. Bars are coloured by trust tier.
- artificialanalysis.ai111
- api.llm-stats.com18
- arcprize.org15
- deepmind.google14
- livebench.ai7
- deploymentsafety.openai.com4
- epoch.ai3
- datasets-server.huggingface.co2
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name GPT-5.6 Terra at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | gpt-5.6 terra (low) | gpt-5-6-terra-low |
| medium | gpt-5.6 terra (medium) | gpt-5-6-terra-medium |
| high | gpt-5.6 terra (high) | gpt-5-6-terra-high |
| xhigh | gpt-5.6 terra (xhigh) | gpt-5-6-terra-xhigh gpt-5.6-terra-xhigh gpt-5.6-terra-xhigh (codex-harness) |
| max | gpt-5.6 terra (max) gpt-5.6 terra max | — |