GPT-5.2 Codex
GPT-5.2 Codex is strong in long context; capable in instruction following, reasoning, and factuality; and behind the leaders in multimodal tasks, coding, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$1.75input$14.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
31results on30benchmarks
- 12 independently verified
- 16 aggregator
- 3 vendor-reported
From 6 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
3 papers reference GPT-5.2 CodexGPT-5.2 Codex benchmark results
31 results on 30 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
7.9% behind the leader1 of 3 ranked benchmarks measured
- 82.33Oct 8, 2026
23.3% behind the leader2 of 3 ranked benchmarks measured
- IFBench77.62Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following66.45Oct 8, 2026livebench_instruction_following@2026-06-25
24.7% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond89.90Oct 8, 2026gpqa
- LiveBench · Reasoning77.71Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam35.73Oct 8, 2026aa_hle
- 8.67Oct 8, 2026
25.0% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy41.07Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination26.64Oct 8, 2026omniscienceNonHallucination
27.6% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro76.30Oct 8, 2026aa_mmmu_pro
28.3% behind the leader7 of 10 ranked benchmarks measured
- LiveBench · Coding83.62Oct 8, 2026livebench_coding@2026-06-25
- SciCode54.63Sep 4, 2026aa_scicode
- 72.80Sep 1, 2026
- LiveBench · Agentic Coding49.39Oct 8, 2026livebench_agentic_coding@2026-06-25
- 56.40Oct 7, 2026
- Terminal-Bench Hard37.12Oct 8, 2026aa_terminalbench_hard
- 1338.68May 22, 2026
Show 2 more coding resultsHide 2 coding results
- 66.30May 1, 2026
- 41.04Oct 8, 2026
35.6% behind the leader1 of 7 ranked benchmarks measured
- 39.38Jun 15, 2026
38.6% behind the leader1 of 5 ranked benchmarks measured
- LiveBench · Mathematics88.77Oct 8, 2026livebench_math@2026-06-25
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language73.68Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis78.20Oct 8, 2026livebench_data_analysis@2026-06-25
- 72.80May 1, 2026
- τ²-Bench Telecom (AA run)92.11Oct 8, 2026aa_tau2
- AA Intelligence28.50Oct 8, 2026aa_intelligence_index
- -2.17Oct 8, 2026
Show 4 more resultsHide 4 results
- 56.52Jun 18, 2026
- 42.96Jun 18, 2026
- 74.30Oct 7, 2026
- 64.00Oct 7, 2026
GPT-5.2 Codex: common questions
Who makes GPT-5.2 Codex?
GPT-5.2 Codex is made by OpenAI.
When was GPT-5.2 Codex released?
GPT-5.2 Codex was released on Dec 11, 2025, according to Artificial Analysis.
What is GPT-5.2 Codex good at?
GPT-5.2 Codex is strong in long context; capable in instruction following, reasoning, and factuality; and behind the leaders in multimodal tasks, coding, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5.2 Codex cost?
GPT-5.2 Codex costs $1.75 per million input tokens and $14.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 86% of the 331 priced models we track.
How many benchmarks has GPT-5.2 Codex been tested on?
We track 31 results for GPT-5.2 Codex on 30 benchmarks from 6 sources, 12 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.2 Codex support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.2 Codex.
About this record
Where GPT-5.2 Codex's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- May 29, 2026
Where the results come from
Verification: 31 scores · 12 independently verified · 16 aggregator-attributed · 3 vendor-reported. How these tiers are assigned
From 6 sources on 6 sites. Artificial Analysis supplies 16 of them; the 12 independently verified results come from 4 sites. Bars are coloured by trust tier.
- artificialanalysis.ai16
- livebench.ai7
- api.llm-stats.com3
- swebench.com3
- datasets-server.huggingface.co1
- labs.scale.com1