GPT-5.3 Codex
GPT-5.3 Codex is strong in long context; capable in reasoning and instruction following; and behind the leaders in coding, multimodal tasks, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$1.75input$14.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
28results on25benchmarks
- 5 independently verified
- 16 aggregator
- 4 vendor-reported
- 3 cross-referenced
From 6 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
4 papers reference GPT-5.3 CodexGPT-5.3 Codex benchmark results
28 results on 25 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
5.9% behind the leader1 of 3 ranked benchmarks measured
- 83.33Oct 8, 2026
17.5% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond91.52Oct 8, 2026gpqa
- Humanity's Last Exam42.49Oct 8, 2026aa_hle
- 16.86Oct 8, 2026
24.5% behind the leader1 of 3 ranked benchmarks measured
- IFBench75.37Oct 8, 2026aa_ifbench
25.4% behind the leader4 of 10 ranked benchmarks measured
- Terminal-Bench Hard53.03Oct 8, 2026aa_terminalbench_hard
- SciCode53.24Sep 4, 2026aa_scicode
- 56.80Oct 7, 2026
- 1408.59May 22, 2026
Show 2 more coding resultsHide 2 coding results
- 56.80May 1, 2026
- 77.30May 1, 2026
26.2% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro78.50Oct 8, 2026aa_mmmu_pro
31.5% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy52.88Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination10.82Oct 8, 2026omniscienceNonHallucination
34.4% behind the leader2 of 7 ranked benchmarks measured
- 64.70Oct 7, 2026
- 48.97Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 4.33Oct 8, 2026
- 99.49May 10, 2026
- 25.00May 10, 2026
- 75.64May 10, 2026
- 10.87Oct 8, 2026
- AA Intelligence32.50Oct 8, 2026aa_intelligence_index
Show 6 more resultsHide 6 results
- 60.54Jun 18, 2026
- 53.10Jun 18, 2026
- 72.76Oct 7, 2026
- 77.30Oct 7, 2026
- 64.70May 6, 2026
- τ²-Bench Telecom (AA run)85.96Oct 8, 2026aa_tau2
GPT-5.3 Codex: common questions
Who makes GPT-5.3 Codex?
GPT-5.3 Codex is made by OpenAI.
When was GPT-5.3 Codex released?
GPT-5.3 Codex was released on Feb 5, 2026, according to Artificial Analysis.
What is GPT-5.3 Codex good at?
GPT-5.3 Codex is strong in long context; capable in reasoning and instruction following; and behind the leaders in coding, multimodal tasks, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does GPT-5.3 Codex cost?
GPT-5.3 Codex costs $1.75 per million input tokens and $14.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 86% of the 331 priced models we track.
How many benchmarks has GPT-5.3 Codex been tested on?
We track 28 results for GPT-5.3 Codex on 25 benchmarks from 6 sources, 5 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.3 Codex support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.3 Codex.
About this record
Where GPT-5.3 Codex's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 29, 2026
Where the results come from
Verification: 28 scores · 5 independently verified · 16 aggregator-attributed · 3 vendor cross-reference · 4 vendor-reported. How these tiers are assigned
From 6 sources on 6 sites. Artificial Analysis supplies 16 of them; the 5 independently verified results come from 3 sites. Bars are coloured by trust tier.
- artificialanalysis.ai16
- api.llm-stats.com4
- deepmind.google3
- raw.githubusercontent.com3
- datasets-server.huggingface.co1
- labs.scale.com1