Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.3 Codex

Score basis

GPT-5.3 Codex is strong in long context; capable in reasoning and instruction following; and behind the leaders in coding, multimodal tasks, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong1 capability
  1. Long Context−5.9%1 of 3
Capable2 capabilities
  1. Reasoning−17.5%3 of 6
  2. Instruction Following−24.5%1 of 3
Limited4 capabilities
  1. Coding−25.4%4 of 10
  2. Multimodal−26.2%1 of 6
  3. Factuality−31.5%2 of 4
  4. Agentic−34.4%2 of 7
Not rated3 capabilities

Too few results yet to rate Safety, Math or Multilingual.

Price

$1.75input$14.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 86% of 331 priced models · 3:1 input-to-output blend, log scale

Evidence

28results on25benchmarks

  • 5 independently verified
  • 16 aggregator
  • 4 vendor-reported
  • 3 cross-referenced

From 6 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

GPT-5.3 Codex benchmark results

28 results on 25 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

StrongFull long context ranking

5.9% behind the leader1 of 3 ranked benchmarks measured

Reasoning

CapableFull reasoning ranking

17.5% behind the leader3 of 6 ranked benchmarks measured

Instruction Following

CapableFull instruction following ranking

24.5% behind the leader1 of 3 ranked benchmarks measured

Coding

LimitedFull coding ranking

25.4% behind the leader4 of 10 ranked benchmarks measured

Show 2 more coding resultsHide 2 coding results

Multimodal

LimitedFull multimodal ranking

26.2% behind the leader1 of 6 ranked benchmarks measured

Factuality

LimitedFull factuality ranking

31.5% behind the leader2 of 4 ranked benchmarks measured

Agentic

LimitedFull agentic ranking

34.4% behind the leader2 of 7 ranked benchmarks measured

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 6 more resultsHide 6 results

GPT-5.3 Codex: common questions

Who makes GPT-5.3 Codex?

GPT-5.3 Codex is made by OpenAI.

When was GPT-5.3 Codex released?

GPT-5.3 Codex was released on Feb 5, 2026, according to Artificial Analysis.

What is GPT-5.3 Codex good at?

GPT-5.3 Codex is strong in long context; capable in reasoning and instruction following; and behind the leaders in coding, multimodal tasks, factuality, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

How much does GPT-5.3 Codex cost?

GPT-5.3 Codex costs $1.75 per million input tokens and $14.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 86% of the 331 priced models we track.

How many benchmarks has GPT-5.3 Codex been tested on?

We track 28 results for GPT-5.3 Codex on 25 benchmarks from 6 sources, 5 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does GPT-5.3 Codex support?

OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.3 Codex.

About this record

Where GPT-5.3 Codex's numbers come from, and every name it appears under.

Tracked since
Apr 27, 2026
Newest source mention
Aug 29, 2026

Where the results come from

Verification: 28 scores · 5 independently verified · 16 aggregator-attributed · 3 vendor cross-reference · 4 vendor-reported. How these tiers are assigned

From 6 sources on 6 sites. Artificial Analysis supplies 16 of them; the 5 independently verified results come from 3 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai16
  • api.llm-stats.com4
  • deepmind.google3
  • raw.githubusercontent.com3
  • datasets-server.huggingface.co1
  • labs.scale.com1

Also known as

gpt-5.3-codex thinking (xhigh)gpt-5-3-codexgpt-5.3 codex (xhigh)gpt codex 5.3-xhighgpt-5.3-codex (codex-harness)