GPT-5.2
GPT-5.2 is strong in long context; capable in reasoning; and behind the leaders in multimodal tasks, math, factuality, coding, instruction following, and agentic tasks. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$1.75input$14.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
204results on110benchmarks
- 64 independently verified
- 48 aggregator
- 25 vendor-reported
- 67 cross-referenced
From 33 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
70 papers reference GPT-5.2GPT-5.2 benchmark results
204 results on 110 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
6.1% behind the leader3 of 3 ranked benchmarks measured
- 82.67Oct 8, 2026
- 83.80May 1, 2026
- 54.50Jun 15, 2026
21.2% behind the leader6 of 6 ranked benchmarks measured
- GPQA Diamond90.30Oct 8, 2026gpqa
- LiveBench · Reasoning83.21Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam37.67Oct 8, 2026aa_hle
- ARC-AGI-252.90Oct 7, 2026ARC-AGI v2
- 45.80May 10, 2026
- 11.55Oct 8, 2026
Show 20 more reasoning resultsHide 20 reasoning results
- 26.67Sep 21, 2026
- 52.91May 10, 2026
- 43.33May 10, 2026
- 9.72May 10, 2026
- 0.83May 10, 2026
- 52.90May 1, 2026
- 7.86Oct 8, 2026
- 0.57Oct 8, 2026
- GPQA Diamond86.36Oct 8, 2026gpqa
- GPQA Diamond71.21Oct 8, 2026gpqa
- GPQA Diamond92.40Oct 7, 2026GPQA
- GPQA Diamond90.00Aug 24, 2026GPQA-D
- 92.40Jun 15, 2026
- Humanity's Last Exam26.69Oct 8, 2026aa_hle
- Humanity's Last Exam8.02Oct 8, 2026aa_hle
- 34.50Oct 7, 2026
- Humanity's Last Exam31.40Aug 24, 2026HLE w/o tools
- Humanity's Last Exam34.50Jun 15, 2026HLE-Full
- 35.40May 17, 2026
- 45.50May 1, 2026
25.7% behind the leader5 of 6 ranked benchmarks measured
- 86.00Jun 15, 2026
- MathVista82.80Jun 15, 2026MathVista (mini)
- CharXiv (reasoning)82.10Oct 7, 2026CharXiv-R
- 80.70Jun 15, 2026
- MMMU-Pro74.57Oct 8, 2026aa_mmmu_pro
Show 9 more multimodal resultsHide 9 multimodal results
- CharXiv (reasoning)82.10Jun 15, 2026CharXiv (RQ)
- 1228.94Sep 22, 2026
- 1245.64Sep 21, 2026
- 1230Jun 17, 2026
- 1245May 25, 2026
- MMMU-Pro65.84Oct 8, 2026aa_mmmu_pro
- 79.50Oct 7, 2026
- 79.50Jun 15, 2026
- 50.50May 22, 2026
31.5% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics93.17Oct 8, 2026livebench_math@2026-06-25
- 67.40Sep 21, 2026
- 31.70Sep 21, 2026
Show 5 more math resultsHide 5 math results
- 98.33Sep 2, 2026
- 98.33May 2, 2026
- 96.97Sep 2, 2026
- 96.97May 10, 2026
- 86.30Jun 15, 2026
31.9% behind the leader4 of 4 ranked benchmarks measured
- Vectara HHEM hallucination ratelower is better10.80Jun 22, 2026
- AA-Omniscience · Accuracy44.32Oct 8, 2026omniscienceAccuracy
- 37.10May 20, 2026
- AA-Omniscience · Non-hallucination18.83Oct 8, 2026omniscienceNonHallucination
Show 7 more factuality resultsHide 7 factuality results
- AA-Omniscience · Accuracy38.28Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy30.85Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination38.40Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination37.67Oct 8, 2026omniscienceNonHallucination
- 34.30May 20, 2026
- 32.70May 20, 2026
- 32.80May 20, 2026
32.3% behind the leader8 of 10 ranked benchmarks measured
- 80.00Oct 7, 2026
- LiveBench · Coding76.07Oct 8, 2026livebench_coding@2026-06-25
- SciCode52.08Sep 4, 2026aa_scicode
- Terminal-Bench Hard46.97Oct 8, 2026aa_terminalbench_hard
- 66.70Sep 25, 2026
- LiveBench · Agentic Coding50.25Oct 8, 2026livebench_agentic_coding@2026-06-25
- 1416.07May 22, 2026
- 29.94Oct 8, 2026
Show 14 more coding resultsHide 14 coding results
- SciCode46.18Sep 4, 2026aa_scicode
- SciCode40.39Sep 4, 2026aa_scicode
- 52.00Aug 24, 2026
- 52.10Jun 15, 2026
- 72.00Jun 15, 2026
- 55.60May 1, 2026
- 71.80Sep 25, 2026
- 72.80Sep 25, 2026
- 69.00Sep 1, 2026
- 80.00Jun 15, 2026
- SWE-bench Verified74.20Jun 5, 2026SWE-bench Verified (mini-swe-agent)
- Terminal-Bench Hard43.18Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard31.82Oct 8, 2026aa_terminalbench_hard
- 62.20May 1, 2026
33.0% behind the leader2 of 3 ranked benchmarks measured
- IFBench75.44Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following61.77Oct 8, 2026livebench_instruction_following@2026-06-25
38.2% behind the leader4 of 7 ranked benchmarks measured
- 65.80Oct 7, 2026
- 60.60Oct 7, 2026
- 48.34Jun 15, 2026
- τ-Bench V3 · Banking11.13Oct 8, 2026tauBanking
Show 7 more agentic resultsHide 7 agentic results
- 23.00May 1, 2026
- 65.80Jun 15, 2026
- 36.19Jun 15, 2026
- 45.17Jun 15, 2026
- 67.60Oct 8, 2026
- MCP Atlas68.00May 17, 2026MCP-Atlas (Public Set)
- 60.60May 1, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis78.16Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language79.81Oct 8, 2026livebench_language@2026-06-25
- 72.67Sep 21, 2026
- 86.53Sep 2, 2026
- frontiermath_tier_4_v118.80Aug 29, 2026frontiermath_tier_4
- 72.80Aug 29, 2026
Show 96 more resultsHide 96 results
- 54.93Jun 18, 2026
- 39.53Jun 18, 2026
- 60.20Jun 18, 2026
- AA Intelligence42.00Jul 3, 2026Artificial Analysis Intelligence Index
- AA Intelligence26.50Oct 8, 2026aa_intelligence_index
- AA Intelligence16.98Oct 8, 2026aa_intelligence_index
- AA Intelligence30.45Oct 8, 2026aa_intelligence_index
- 0.27Oct 8, 2026
- -0.88Oct 8, 2026
- -12.25Oct 8, 2026
- 100.00May 20, 2026
- 100.00Oct 7, 2026
- 100.00Jun 15, 2026
- AIME 202598.00Jun 5, 2026AIME25
- AIME25 no tools98.00Aug 24, 2026AIME25
- 86.17May 10, 2026
- 78.67May 10, 2026
- 55.67May 10, 2026
- 12.33May 10, 2026
- 1437Jul 23, 2026
- 34.68Jun 18, 2026
- 44.18Jun 18, 2026
- 48.67Jun 18, 2026
- BrowseComp (context management)57.80Aug 24, 2026BrowseComp (w/ctx manage)
- 70.00Jun 5, 2026
- browsecomp_with_context_manager65.80May 17, 2026BrowseComp (w/ Context Manage)
- 76.10May 17, 2026
- DeepSearchQA (F1)71.30Jun 15, 2026DeepSearchQA
- FrontierMath (overall)40.30Oct 7, 2026FrontierMath
- frontiermath_tier_4_v116.70May 20, 2026frontiermath_tier_4
- frontiermath_tier_4_v16.25May 20, 2026frontiermath_tier_4
- 92.47May 22, 2026
- 74.84May 22, 2026
- 57.14May 22, 2026
- 77.41May 22, 2026
- 1462.00May 1, 2026
- 92.40May 1, 2026
- 94.40Jul 9, 2026
- 34.30Jul 9, 2026
- 56.80Jul 9, 2026
- 45.90Jul 9, 2026
- HLE (with tools)45.50Jun 15, 2026HLE-Full (w/ tools)
- 99.40Oct 7, 2026
- 100.00Jun 23, 2026
- HMMT Feb. 202599.40Jun 15, 2026HMMT 2025 (Feb)
- 95.83May 10, 2026
- 99.17May 10, 2026
- 97.10May 17, 2026
- Image input eval - extremism (not_unsafe)0.99Jul 9, 2026
- Image input eval - harms-erotic (not_unsafe)1.00Jul 9, 2026
- Image input eval - hate (not_unsafe)0.99Jul 9, 2026
- Image input eval - self-harm (not_unsafe)0.99Jul 9, 2026
- 84.00Jun 15, 2026
- 74.84Oct 7, 2026
- 89.00Jun 5, 2026
- 2393.00May 1, 2026
- 76.50Jun 15, 2026
- 83.00Jun 15, 2026
- 76.96May 10, 2026
- 35.00May 10, 2026
- 99.62May 10, 2026
- 86.70Jun 15, 2026
- 87.00Jun 5, 2026
- 89.60Oct 7, 2026
- MMMLU89.60May 1, 2026mmmlu_multilingual_qa
- 80.80Jun 15, 2026
- 64.80Jun 15, 2026
- 85.70Jun 15, 2026
- 63.70Jun 15, 2026
- ScreenSpot-Pro (No tools)86.30Oct 7, 2026ScreenSpot Pro
- 45.00Aug 24, 2026
- 55.80Jun 15, 2026
- 69.00Aug 10, 2026
- 71.80May 1, 2026
- 3.60Jun 5, 2026
- 80.70Jun 5, 2026
- 54.00Jun 15, 2026
- 46.30May 17, 2026
- 46.30Oct 7, 2026
- 41.70Jun 5, 2026
- 100.00Jun 22, 2026
- 186.30Jun 22, 2026
- 89.20Jun 22, 2026
- 85.90Oct 7, 2026
- 85.90Jun 15, 2026
- 28.00Jun 15, 2026
- 9.00Jun 15, 2026
- 7.00Jun 15, 2026
- 85.50Jun 13, 2026
- τ²-Bench85.00Jun 5, 2026𝜏²-Bench Telecom
- 82.00May 6, 2026
- τ²-Bench (Retail)82.00Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)74.27Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)46.49Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)84.80Oct 8, 2026aa_tau2
- 98.70May 1, 2026
GPT-5.2: common questions
Who makes GPT-5.2?
GPT-5.2 is made by OpenAI.
When was GPT-5.2 released?
GPT-5.2 was released on Dec 11, 2025, according to Artificial Analysis.
What is GPT-5.2 good at?
GPT-5.2 is strong in long context; capable in reasoning; and behind the leaders in multimodal tasks, math, factuality, coding, instruction following, and agentic tasks. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5.2 cost?
GPT-5.2 costs $1.75 per million input tokens and $14.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 86% of the 330 priced models we track.
How many benchmarks has GPT-5.2 been tested on?
We track 204 results for GPT-5.2 on 110 benchmarks from 33 sources, 64 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.2 support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.2.
About this record
Where GPT-5.2's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Aug 23, 2026
Where the results come from
Verification: 204 scores · 64 independently verified · 48 aggregator-attributed · 67 vendor cross-reference · 25 vendor-reported. How these tiers are assigned
From 33 sources on 16 sites. Hugging Face supplies 58 of them; the 64 independently verified results come from 13 sites. Bars are coloured by trust tier.
- huggingface.co58
- artificialanalysis.ai49
- api.llm-stats.com17
- deepmind.google13
- arcprize.org10
- epoch.ai9
- matharena.ai9
- deploymentsafety.openai.com8
- livebench.ai7
- raw.githubusercontent.com7
- swebench.com7
- datasets-server.huggingface.co3
- lmarena.ai3
- labs.scale.com2
- 99franklin.github.io1
- simple-bench.com1
Also known as
How our sources name GPT-5.2 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | — | gpt-5.2 (low) gpt-5.2-low-2025-12-11 |
| medium | — | gpt-5-2-medium gpt-5.2 (medium) |
| high | GPT 5.2 (2025-12-11) (high) gpt 5.2 (high) | gpt-5-2 (high reasoning) gpt-5.2 (2025-12-11) (high reasoning) gpt-5.2 (high reasoning) gpt-5.2-high gpt-5.2-high-2025-12-11 |
| xhigh | GPT-5.2 Thinking (xhigh) | gpt-5.2 (xhigh) GPT-5.2 (xhigh) (Non-Reasoning) gpt-5.2 (xhigh) (reasoning) |