Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.2

Score basis

GPT-5.2 is strong in long context; capable in reasoning; and behind the leaders in multimodal tasks, math, factuality, coding, instruction following, and agentic tasks. Too few results yet to rate safety or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong1 capability
  1. Long Context−6.1%3 of 3
Capable1 capability
  1. Reasoning−21.2%6 of 6
Limited6 capabilities
  1. Multimodal−25.7%5 of 6
  2. Math−31.5%3 of 5
  3. Factuality−31.9%4 of 4
  4. Coding−32.3%8 of 10
  5. Instruction Following−33.0%2 of 3
  6. Agentic−38.2%4 of 7
Not rated2 capabilities

Too few results yet to rate Safety or Multilingual.

Price

$1.75input$14.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 86% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

204results on110benchmarks

  • 64 independently verified
  • 48 aggregator
  • 25 vendor-reported
  • 67 cross-referenced

From 33 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

GPT-5.2 benchmark results

204 results on 110 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

StrongFull long context ranking

6.1% behind the leader3 of 3 ranked benchmarks measured

Show 2 more long context resultsHide 2 long context results

Reasoning

CapableFull reasoning ranking

21.2% behind the leader6 of 6 ranked benchmarks measured

Show 20 more reasoning resultsHide 20 reasoning results

Multimodal

LimitedFull multimodal ranking

25.7% behind the leader5 of 6 ranked benchmarks measured

Show 9 more multimodal resultsHide 9 multimodal results

31.5% behind the leader3 of 5 ranked benchmarks measured

Show 5 more math resultsHide 5 math results

Factuality

LimitedFull factuality ranking

31.9% behind the leader4 of 4 ranked benchmarks measured

Show 7 more factuality resultsHide 7 factuality results

Coding

LimitedFull coding ranking

32.3% behind the leader8 of 10 ranked benchmarks measured

Show 14 more coding resultsHide 14 coding results

Instruction Following

LimitedFull instruction following ranking

33.0% behind the leader2 of 3 ranked benchmarks measured

Show 3 more instruction following resultsHide 3 instruction following results

Agentic

LimitedFull agentic ranking

38.2% behind the leader4 of 7 ranked benchmarks measured

Show 7 more agentic resultsHide 7 agentic results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 96 more resultsHide 96 results

GPT-5.2: common questions

Who makes GPT-5.2?

GPT-5.2 is made by OpenAI.

When was GPT-5.2 released?

GPT-5.2 was released on Dec 11, 2025, according to Artificial Analysis.

What is GPT-5.2 good at?

GPT-5.2 is strong in long context; capable in reasoning; and behind the leaders in multimodal tasks, math, factuality, coding, instruction following, and agentic tasks. Too few results yet to rate safety or multilingual tasks.

How much does GPT-5.2 cost?

GPT-5.2 costs $1.75 per million input tokens and $14.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 86% of the 330 priced models we track.

How many benchmarks has GPT-5.2 been tested on?

We track 204 results for GPT-5.2 on 110 benchmarks from 33 sources, 64 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does GPT-5.2 support?

OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.2.

About this record

Where GPT-5.2's numbers come from, and every name it appears under.

Tracked since
May 2, 2026
Newest source mention
Aug 23, 2026

Where the results come from

Verification: 204 scores · 64 independently verified · 48 aggregator-attributed · 67 vendor cross-reference · 25 vendor-reported. How these tiers are assigned

From 33 sources on 16 sites. Hugging Face supplies 58 of them; the 64 independently verified results come from 13 sites. Bars are coloured by trust tier.

  • huggingface.co58
  • artificialanalysis.ai49
  • api.llm-stats.com17
  • deepmind.google13
  • arcprize.org10
  • epoch.ai9
  • matharena.ai9
  • deploymentsafety.openai.com8
  • livebench.ai7
  • raw.githubusercontent.com7
  • swebench.com7
  • datasets-server.huggingface.co3
  • lmarena.ai3
  • labs.scale.com2
  • 99franklin.github.io1
  • simple-bench.com1

Also known as

How our sources name GPT-5.2 at each reasoning setting.

SettingShort formAPI id
low—gpt-5.2 (low) gpt-5.2-low-2025-12-11
medium—gpt-5-2-medium gpt-5.2 (medium)
highGPT 5.2 (2025-12-11) (high) gpt 5.2 (high)gpt-5-2 (high reasoning) gpt-5.2 (2025-12-11) (high reasoning) gpt-5.2 (high reasoning) gpt-5.2-high gpt-5.2-high-2025-12-11
xhighGPT-5.2 Thinking (xhigh)gpt-5.2 (xhigh) GPT-5.2 (xhigh) (Non-Reasoning) gpt-5.2 (xhigh) (reasoning)
Also listed asgpt-5-2gpt-5-2-non-reasoninggpt-5.2 (2025-12-11)gpt-5.2 (non-reasoning)gpt-5.2 (thinking)gpt-5.2-20251211gpt-5.2-thinkinggpt-5.2:batchgpt5.2gpt-5.2-2025-12-11