GPT-5.6 Sol
GPT-5.6 Sol is at the frontier in agentic tasks; strong in reasoning, long context, coding, and math; and capable in multimodal tasks, factuality, and instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$4.00input$20.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
199results on119benchmarks
- 28 independently verified
- 70 aggregator
- 27 vendor-reported
- 74 cross-referenced
From 33 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
113 papers reference GPT-5.6 SolGPT-5.6 Sol benchmark results
199 results on 119 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
4.2% behind the leader7 of 7 ranked benchmarks measured
- 90.40Oct 7, 2026
- 83.00Sep 22, 2026
- 81.80Oct 8, 2026
- τ-Bench V3 · Banking44.33Oct 8, 2026tauBanking
- 55.57Oct 8, 2026
- Terminal-Bench 4.039.90Oct 8, 2026
Show 15 more agentic resultsHide 15 agentic results
- AA ApexAgents56.70Aug 24, 2026APEX-Agents
- AA ApexAgents39.90Jul 27, 2026APEX-Agents
- 56.21Oct 8, 2026
- 90.40Jul 27, 2026
- 53.62Oct 8, 2026
- 50.23Oct 8, 2026
- 37.12Oct 8, 2026
- 83.60Jul 27, 2026
- 81.80Jul 16, 2026
- Terminal-Bench 4.024.75Oct 8, 2026
- Terminal-Bench 4.020.71Oct 8, 2026
- Terminal-Bench 4.014.65Oct 8, 2026
- τ-Bench V3 · Banking38.14Oct 8, 2026tauBanking
- τ-Bench V3 · Banking36.70Oct 8, 2026tauBanking
- τ-Bench V3 · Banking19.59Oct 8, 2026tauBanking
2.6% behind the leader6 of 6 ranked benchmarks measured
- 32.29Oct 8, 2026
- LiveBench · Reasoning91.65Oct 8, 2026livebench_reasoning@2026-06-25
- GPQA Diamond94.14Oct 8, 2026gpqa
- 92.50Sep 21, 2026
- Humanity's Last Exam49.49Oct 8, 2026aa_hle
- 64.80Sep 21, 2026
Show 14 more reasoning resultsHide 14 reasoning results
- 90.00Sep 21, 2026
- 85.42Sep 21, 2026
- 7.78Sep 21, 2026
- 6.99Sep 21, 2026
- 7.78Oct 7, 2026
- 28.57Oct 8, 2026
- 25.71Oct 8, 2026
- 5.14Oct 8, 2026
- GPQA Diamond93.13Oct 8, 2026gpqa
- GPQA Diamond92.83Oct 8, 2026gpqa
- GPQA Diamond92.63Oct 8, 2026gpqa
- Humanity's Last Exam47.31Oct 8, 2026aa_hle
- Humanity's Last Exam46.01Oct 8, 2026aa_hle
- Humanity's Last Exam16.68Oct 8, 2026aa_hle
5.1% behind the leader1 of 3 ranked benchmarks measured
- 84.00Oct 8, 2026
Show 4 more long context resultsHide 4 long context results
- 82.33Oct 8, 2026
- 81.67Oct 8, 2026
- 62.33Oct 8, 2026
- 67.10Aug 24, 2026
6.2% behind the leader7 of 10 ranked benchmarks measured
- Terminal-Bench Hard65.91Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.188.01Oct 8, 2026terminalbenchV21
- LiveBench · Coding83.94Oct 8, 2026livebench_coding@2026-06-25
- 1617.76Sep 21, 2026
- SciCode57.06Oct 8, 2026aa_scicode
- LiveBench · Agentic Coding56.21Oct 8, 2026livebench_agentic_coding@2026-06-25
- 64.60Oct 7, 2026
Show 8 more coding resultsHide 8 coding results
- SciCode57.75Oct 8, 2026aa_scicode
- SciCode57.41Oct 8, 2026aa_scicode
- 64.60Sep 1, 2026
- Terminal-Bench 2.189.51Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.187.27Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.174.16Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard61.36Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard62.12Oct 8, 2026aa_terminalbench_hard
9.3% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics96.20Oct 8, 2026livebench_math@2026-06-25
- 89.12Sep 21, 2026
- FrontierMath Tier 483.00Oct 7, 2026FrontierMath Tier 4 (v2)
Show 2 more math resultsHide 2 math results
- 99.90Jul 16, 2026
- 82.93Sep 21, 2026
18.3% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro83.41Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)85.80Sep 11, 2026CharXiv Reasoning (No tools)
Show 5 more multimodal resultsHide 5 multimodal results
- CharXiv (reasoning)84.60Aug 24, 2026CharXiv (RQ)
- 1279.21Sep 21, 2026
- MMMU-Pro82.66Oct 8, 2026aa_mmmu_pro
- MMMU-Pro81.85Oct 8, 2026aa_mmmu_pro
- MMMU-Pro71.91Oct 8, 2026aa_mmmu_pro
18.7% behind the leader3 of 4 ranked benchmarks measured
- 69.70Sep 21, 2026
- AA-Omniscience · Accuracy59.40Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination7.80Oct 8, 2026omniscienceNonHallucination
Show 8 more factuality resultsHide 8 factuality results
- AA-Omniscience · Accuracy58.82Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy58.35Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy48.70Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination8.13Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination8.80Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination9.23Oct 8, 2026omniscienceNonHallucination
- 71.60Jul 16, 2026
- 12.40Sep 23, 2026
21.0% behind the leader2 of 3 ranked benchmarks measured
- LiveBench · Instruction Following71.85Oct 8, 2026livebench_instruction_following@2026-06-25
- IFBench72.65Oct 8, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 1484Oct 8, 2026
- livebench_data_analysis79.84Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language87.68Oct 8, 2026livebench_language@2026-06-25
- 32.33Oct 8, 2026
- vectara_avg_summary_length131.10Sep 23, 2026Average Summary Length (Words)
- vectara_answer_rate99.10Sep 23, 2026Answer Rate
Show 105 more resultsHide 105 results
- 50.50Sep 9, 2026
- 47.82Sep 9, 2026
- 44.70Sep 9, 2026
- AA Intelligence46.97Oct 8, 2026aa_intelligence_index
- AA Intelligence44.01Oct 8, 2026aa_intelligence_index
- AA Intelligence42.35Oct 8, 2026aa_intelligence_index
- 21.97Oct 8, 2026
- 20.98Oct 8, 2026
- 20.37Oct 8, 2026
- 19.40Oct 8, 2026
- 52.70Oct 7, 2026
- 30.80Sep 22, 2026
- 28.60Aug 28, 2026
- 53.60Aug 12, 2026
- 74.00Aug 12, 2026
- 96.50Sep 21, 2026
- 97.50Sep 21, 2026
- 97.00Sep 21, 2026
- 59.00Oct 7, 2026
- Artificial Analysis Coding Index78.35Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index77.39Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index77.16Sep 9, 2026aa_coding_index
- baby_vision_with_python88.90Aug 24, 2026BabyVision w/ python
- 89.40Jul 16, 2026
- 79.90Sep 10, 2026
- charxiv_rq_with_python87.80Jul 16, 2026Charxiv RQ (with python)
- CyberGym84.50Sep 10, 2026CyberGym (Pass@1)
- 83.60Aug 28, 2026
- 72.70Oct 7, 2026
- DeepSWE72.70Aug 28, 2026DeepSWE (v1.1)
- 73.00Oct 7, 2026
- DeepSWE 1.173.00Sep 22, 2026DeepSWE v1.1
- DeepSWE v1.1 (Resolved)73.00Sep 10, 2026
- 13.70Jul 7, 2026
- 53.80Aug 24, 2026
- FrontierMath (overall)89.00Oct 7, 2026FrontierMath
- GDP (Surge AI)30.70Oct 7, 2026GDP.pdf
- GDP.PDF (All pass rate)40.00Sep 9, 2026
- GDPval-AA 2.11588.00Sep 22, 2026GDPval-AA v2.1
- 1711.00Sep 10, 2026
- GDPval-AA v2 Elo1710.00Sep 11, 2026GDPVal-AA v2 (Elo)
- GDPval-AA v2 Elo1736.00Jul 27, 2026GDPval-AA v2 (Elo)
- 91.80Jul 16, 2026
- 7.60Aug 11, 2026
- 57.00Oct 7, 2026
- 55.30Aug 12, 2026
- 95.50Sep 29, 2026
- 33.10Sep 29, 2026
- 32.00Jul 9, 2026
- 57.00Sep 29, 2026
- 60.50Oct 7, 2026
- 60.50Jul 24, 2026
- 60.50Sep 29, 2026
- HLE (with tools)64.50Aug 28, 2026HLE w/ Tools
- HLE (with tools)58.00Aug 12, 2026HLE w/ tools
- 54.50Sep 9, 2026
- Image input eval - extremism (not_unsafe)0.97Jul 9, 2026
- Image input eval - harms-erotic (not_unsafe)0.99Jul 9, 2026
- Image input eval - hate (not_unsafe)1.00Jul 9, 2026
- Image input eval - self-harm (not_unsafe)0.99Jul 9, 2026
- 45.40Sep 22, 2026
- 82.10Sep 9, 2026
- 95.80Aug 24, 2026
- 65.60Sep 10, 2026
- 92.90Jul 27, 2026
- MLS Bench Litelower is better46.20Aug 12, 2026
- 81.20Aug 24, 2026
- 93.80Aug 24, 2026
- 55.50Jul 7, 2026
- NL2Repo-Bench (Score)56.80Sep 10, 2026
- 63.20Jul 27, 2026
- 85.80Aug 24, 2026
- 62.60Oct 7, 2026
- 62.60Aug 24, 2026
- OSWorld 2.0 (partial)62.60Sep 9, 2026OSWorld-2.0 (Partial score)
- 90.50Aug 12, 2026
- 36.20Aug 28, 2026
- 34.60Jul 27, 2026
- 25.00Sep 22, 2026
- 77.60Jul 27, 2026
- 23.00Aug 28, 2026
- ProgramBench (Almost@1)23.00Sep 10, 2026
- 35.50Sep 1, 2026
- 73.80Jul 27, 2026
- 32.40Jul 27, 2026
- 98.50Jul 16, 2026
- SWE-Marathon42.50Aug 28, 2026SWE-Marathon (v1.1)
- 39.00Jul 27, 2026
- SWEBench Pro Public64.60Jul 16, 2026SWEBench Pro (Public)
- Terminal-bench 3.034.40Sep 10, 2026Terminal-Bench 3.0 (Pass@1)
- 34.60Aug 28, 2026
- Terminal-Bench 4.037.30Oct 7, 2026
- Terminal-Bench 4.037.30Sep 22, 2026
- Terminal-Bench 4.039.90Sep 22, 2026
- Terminal-Bench-Science 0.122.40Sep 22, 2026
- 58.00Oct 7, 2026
- 74.90Sep 22, 2026
- vectara_factual_consistency87.60Sep 23, 2026Factual Consistency Rate
- VideoMME (w sub.)89.50Aug 24, 2026Video-MME (w. sub)
- ZeroBench17.00Aug 24, 2026ZeroBench (pass@5)
- ZeroBench-main w/ tools (Pass@5)53.00Sep 10, 2026
- τ²-Bench Telecom (AA run)84.80Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)85.09Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)83.33Oct 8, 2026aa_tau2
- τ³-Bench Banking33.00Aug 24, 2026τ³-Banking
GPT-5.6 Sol: common questions
Who makes GPT-5.6 Sol?
GPT-5.6 Sol is made by OpenAI.
When was GPT-5.6 Sol released?
GPT-5.6 Sol was released on Jul 9, 2026, according to Artificial Analysis.
What is GPT-5.6 Sol good at?
GPT-5.6 Sol is at the frontier in agentic tasks; strong in reasoning, long context, coding, and math; and capable in multimodal tasks, factuality, and instruction following. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5.6 Sol cost?
GPT-5.6 Sol costs $4.00 per million input tokens and $20.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 92% of the 330 priced models we track.
How many benchmarks has GPT-5.6 Sol been tested on?
We track 199 results for GPT-5.6 Sol on 119 benchmarks from 33 sources, 28 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.6 Sol support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.6 Sol.
About this record
Where GPT-5.6 Sol's numbers come from, and every name it appears under.
- Tracked since
- Jul 1, 2026
- Newest source mention
- Oct 7, 2026
Where the results come from
Verification: 199 scores · 28 independently verified · 70 aggregator-attributed · 74 vendor cross-reference · 27 vendor-reported. How these tiers are assigned
From 33 sources on 16 sites. Artificial Analysis supplies 70 of them; the 28 independently verified results come from 8 sites. Bars are coloured by trust tier.
- artificialanalysis.ai70
- huggingface.co60
- api.llm-stats.com15
- deploymentsafety.openai.com12
- arcprize.org8
- livebench.ai7
- deepmind.google6
- anthropic.com4
- raw.githubusercontent.com4
- epoch.ai3
- www-cdn.anthropic.com3
- datasets-server.huggingface.co2
- labs.scale.com2
- lmarena.ai1
- simple-bench.com1
- x.ai1
Also known as
How our sources name GPT-5.6 Sol at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | gpt-5.6 sol (low) gpt 5.6 sol low | gpt-5-6-sol-low |
| medium | gpt-5.6 sol (medium) | gpt-5-6-sol-medium |
| high | gpt-5.6 sol (high) gpt-5.6 sol high | gpt-5-6-sol-high |
| xhigh | gpt-5.6 sol (xhigh) | gpt-5.6-sol-xhigh gpt-5.6-sol-xhigh (codex-harness) gpt-5-6-sol-xhigh |
| max | gpt-5.6 sol (max) gpt-5.6 sol max gpt sol 5.6 max gpt-5.6 sol (pro, max) | gpt-5.6-sol (max, via codex) |