GPT-5.5
GPT-5.5 is strong in long context and reasoning; and capable in coding, math, agentic tasks, factuality, multimodal tasks, and instruction following. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$5.00input$30.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
169results on130benchmarks
- 29 independently verified
- 40 aggregator
- 30 vendor-reported
- 70 cross-referenced
From 34 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
224 papers reference GPT-5.5GPT-5.5 benchmark results
169 results on 130 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
4.8% behind the leader1 of 3 ranked benchmarks measured
- 84.33Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 81.00Oct 8, 2026
6.0% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond93.23Oct 8, 2026gpqa
- LiveBench · Reasoning89.65Oct 8, 2026livebench_reasoning@2026-06-25
- ARC-AGI-285.00Oct 7, 2026ARC-AGI v2
- 25.43Oct 8, 2026
- 69.00May 10, 2026
- Humanity's Last Exam45.04Oct 8, 2026aa_hle
Show 5 more reasoning resultsHide 5 reasoning results
- 85.00Sep 28, 2026
- 0.43May 10, 2026
- 8.00Oct 8, 2026
- GPQA Diamond91.01Oct 8, 2026gpqa
- Humanity's Last Exam32.67Oct 8, 2026aa_hle
13.0% behind the leader7 of 10 ranked benchmarks measured
- Terminal-Bench 2.179.40Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard59.85Oct 8, 2026aa_terminalbench_hard
- LiveBench · Coding82.15Oct 8, 2026livebench_coding@2026-06-25
- SciCode56.13Oct 8, 2026aa_scicode
- 1512.71Sep 21, 2026
- LiveBench · Agentic Coding53.99Oct 8, 2026livebench_agentic_coding@2026-06-25
- 58.60Oct 7, 2026
Show 7 more coding resultsHide 7 coding results
- 1456.06Sep 22, 2026
- SciCode54.51Oct 8, 2026aa_scicode
- 58.60Aug 12, 2026
- Terminal-Bench 2.165.54Oct 8, 2026terminalbenchV21
- Terminal-Bench 2.183.40Jul 13, 2026Terminal Bench 2.1 (Best Reported Harness)
- Terminal-Bench 2.184.00Jul 13, 2026Terminal Bench 2.1 (Terminus-2)
- Terminal-Bench Hard52.27Oct 8, 2026aa_terminalbench_hard
13.2% behind the leader5 of 5 ranked benchmarks measured
- LiveBench · Mathematics95.86Oct 8, 2026livebench_math@2026-06-25
- 85.26Sep 21, 2026
- 72.50Sep 21, 2026
Show 5 more math resultsHide 5 math results
- 100.00Sep 2, 2026
- 98.30Jul 13, 2026
- 98.48Sep 2, 2026
- HMMT Feb 202696.70Jul 13, 2026HMMT Feb. 2026
- 98.21Sep 2, 2026
13.5% behind the leader7 of 7 ranked benchmarks measured
- 78.70Oct 7, 2026
- 84.40Oct 7, 2026
- 75.30Oct 7, 2026
- τ-Bench V3 · Banking36.70Oct 8, 2026tauBanking
- 41.46Oct 8, 2026
- Terminal-Bench 4.09.09Oct 8, 2026
Show 9 more agentic resultsHide 9 agentic results
- 37.68Oct 8, 2026
- AA ApexAgents35.40Sep 28, 2026APEX Agents
- 45.81Oct 8, 2026
- 84.40Jul 27, 2026
- 27.44Oct 8, 2026
- 75.30Oct 8, 2026
- 78.70Aug 12, 2026
- Terminal-Bench 4.05.05Oct 8, 2026
- τ-Bench V3 · Banking24.95Oct 8, 2026tauBanking
18.0% behind the leader4 of 4 ranked benchmarks measured
- 9.30May 2, 2026
- AA-Omniscience · Accuracy57.03Oct 8, 2026omniscienceAccuracy
- 63.00Sep 21, 2026
- AA-Omniscience · Non-hallucination10.90Oct 8, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy54.85Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination11.96Oct 8, 2026omniscienceNonHallucination
19.9% behind the leader2 of 6 ranked benchmarks measured
- 84.20Sep 28, 2026
- MMMU-Pro79.02Oct 8, 2026aa_mmmu_pro
Show 5 more multimodal resultsHide 5 multimodal results
- 78.30Sep 28, 2026
- 1292.12Sep 21, 2026
- 1297.29Aug 25, 2026
- MMMU-Pro81.10Oct 8, 2026aa_mmmu_pro
- 61.10Sep 28, 2026
20.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench71.63Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following70.73Oct 8, 2026livebench_instruction_following@2026-06-25
Show 1 more instruction following resultHide 1 instruction following result
- IFBench64.35Oct 8, 2026aa_ifbench
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language87.36Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis81.58Oct 8, 2026livebench_data_analysis@2026-06-25
- 39.67Oct 8, 2026
- 1477Oct 6, 2026
- 95.00Sep 21, 2026
- 94.93Sep 2, 2026
Show 97 more resultsHide 97 results
- 35.36Sep 9, 2026
- 30.62Sep 9, 2026
- AA Intelligence30.71Oct 8, 2026aa_intelligence_index
- AA Intelligence36.98Oct 8, 2026aa_intelligence_index
- 18.75Oct 8, 2026
- 15.10Oct 8, 2026
- 26.60Jul 27, 2026
- 86.00Jul 13, 2026
- Artificial Analysis Coding Index60.90Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index71.57Sep 9, 2026aa_coding_index
- 55.90Sep 28, 2026
- 84.40Aug 12, 2026
- 69.40Sep 28, 2026
- 81.80Oct 7, 2026
- 81.80Sep 28, 2026
- 70.00Sep 28, 2026
- 67.00Jul 27, 2026
- 67.00Oct 7, 2026
- 13.82Jul 7, 2026
- 81.70Sep 28, 2026
- 75.90Sep 28, 2026
- 81.90Sep 28, 2026
- 64.50Sep 28, 2026
- 60.00Oct 7, 2026
- Finance Agent v1.165.30Sep 28, 2026
- 51.76Oct 7, 2026
- 64.30Jul 30, 2026
- FrontierMath (overall)35.40Oct 7, 2026FrontierMath
- frontiermath_tier_4_v135.40Aug 29, 2026frontiermath_tier_4
- GDP (Surge AI)24.90Jun 12, 2026GDP.pdf
- GDPval-AA (Elo)1769.00Jun 12, 2026GDPval-AA
- 1509.00Aug 12, 2026
- 1492.00Jun 30, 2026
- GDPval-AA v2 Elo1491.00Jul 27, 2026GDPval-AA v2 (Elo)
- 25.00Oct 7, 2026
- 20.40Jun 13, 2026
- GraphWalks — Parents task, 256K-token context subset90.10Jun 12, 2026GraphWalks Parents 256K subset
- 45.40Jun 12, 2026
- GraphWalks BFS 256K73.70Jun 12, 2026GraphWalks BFS 256K subset
- 58.50Jun 12, 2026
- 1.50Aug 11, 2026
- Hard-negative protein binding prediction0.40Jul 7, 2026Hard-negative protein binding prediction (pass@4)
- 56.50Jun 12, 2026
- 95.60Sep 29, 2026
- 31.50Jul 9, 2026
- 56.50Sep 10, 2026
- 51.80Aug 12, 2026
- 51.80Sep 29, 2026
- HLE (with tools)52.20Aug 12, 2026Humanity's Last Exam (With tools)
- 44.20Jul 13, 2026
- HLE-Verified (no tool)50.40Sep 28, 2026
- 96.50Jul 13, 2026
- Image input eval - extremism (not_unsafe)0.99Jul 9, 2026
- Image input eval - harms-erotic (not_unsafe)0.99Jul 9, 2026
- Image input eval - hate (not_unsafe)1.00Jul 9, 2026
- Image input eval - self-harm (not_unsafe)0.98Jul 9, 2026
- 38.30Jul 27, 2026
- 55.80Jun 13, 2026
- 2.10Oct 7, 2026
- 2.10Jul 7, 2026
- 80.71Oct 7, 2026
- MathVerse (Vision-Only)84.60Sep 28, 2026
- 92.90Jul 27, 2026
- 25.10Jun 13, 2026
- MLS Bench Litelower is better35.50Jul 27, 2026
- MMSIBench (circular)36.00Sep 28, 2026
- 50.70Jul 13, 2026
- Office QA Pro [Multimodal]69.50Sep 28, 2026
- 54.10Oct 7, 2026
- 62.90Sep 28, 2026
- 6.70Jul 13, 2026
- 28.40Jul 27, 2026
- 70.80Jul 27, 2026
- 0.12Jul 30, 2026
- 82.20Sep 28, 2026
- 64.00Jul 27, 2026
- 58.60Sep 28, 2026
- 29.10Jul 27, 2026
- 72.70Sep 28, 2026
- 14.00Jul 27, 2026
- 12.00Jul 13, 2026
- SWE-Pro Bench58.60Sep 28, 2026
- 82.70Oct 7, 2026
- 55.60Jul 13, 2026
- 55.60Oct 7, 2026
- 55.60Sep 28, 2026
- 73.50Jul 27, 2026
- vectara_answer_rate100.00May 2, 2026Answer Rate
- vectara_avg_summary_length129.60May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency90.70May 2, 2026Factual Consistency Rate
- 43.50Sep 28, 2026
- WMDP85.50Jul 13, 2026WMDP-Chem
- 34.60Sep 28, 2026
- ZeroBench (main)13.00Sep 28, 2026
- ZeroBench (sub)41.00Sep 28, 2026
- τ²-Bench Telecom (AA run)92.98Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)83.92Oct 8, 2026aa_tau2
GPT-5.5: common questions
Who makes GPT-5.5?
GPT-5.5 is made by OpenAI.
When was GPT-5.5 released?
GPT-5.5 was released on Apr 23, 2026, according to Artificial Analysis.
What is GPT-5.5 good at?
GPT-5.5 is strong in long context and reasoning; and capable in coding, math, agentic tasks, factuality, multimodal tasks, and instruction following. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5.5 cost?
GPT-5.5 costs $5.00 per million input tokens and $30.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 94% of the 330 priced models we track.
How many benchmarks has GPT-5.5 been tested on?
We track 169 results for GPT-5.5 on 130 benchmarks from 34 sources, 29 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.5.
About this record
Where GPT-5.5's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 24, 2026
Where the results come from
Verification: 169 scores · 29 independently verified · 40 aggregator-attributed · 70 vendor cross-reference · 30 vendor-reported. How these tiers are assigned
From 34 sources on 18 sites. Artificial Analysis supplies 40 of them; the 29 independently verified results come from 9 sites. Bars are coloured by trust tier.
- artificialanalysis.ai40
- lf3-static.bytednsdoc.com28
- huggingface.co21
- api.llm-stats.com16
- www-cdn.anthropic.com15
- deploymentsafety.openai.com11
- livebench.ai7
- ai.meta.com4
- datasets-server.huggingface.co4
- epoch.ai4
- matharena.ai4
- raw.githubusercontent.com4
- openai.com3
- arcprize.org2
- labs.scale.com2
- thinkingmachines.ai2
- lmarena.ai1
- simple-bench.com1
Also known as
How our sources name GPT-5.5 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | — | gpt-5-5-low gpt-5.5 (low) |
| medium | — | gpt-5-5-medium gpt-5.5 (medium) |
| high | — | gpt-5-5-high gpt-5.5 (high) gpt-5.5-high gpt-5.5-high (codex-harness) |
| xhigh | gpt-5.5 xhigh | gpt-5.5 (xhigh) gpt-5.5-xhigh (codex-harness) |