GPT-5 Codex
GPT-5 Codex is capable in long context; and behind the leaders in instruction following, factuality, reasoning, coding, multimodal tasks, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$1.25input$10.00outputper million tokens
From Artificial Analysis · All prices
Evidence
17results on17benchmarks
- 16 aggregator
- 1 vendor-reported
From 2 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
GPT-5 Codex benchmark results
17 results on 17 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
23.5% behind the leader1 of 3 ranked benchmarks measured
- 71.67Oct 7, 2026
25.2% behind the leader1 of 3 ranked benchmarks measured
- IFBench74.15Oct 7, 2026aa_ifbench
26.2% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy38.70Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination24.58Oct 7, 2026omniscienceNonHallucination
28.1% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond83.74Oct 7, 2026gpqa
- Humanity's Last Exam27.85Oct 7, 2026aa_hle
- 5.14Oct 7, 2026
28.8% behind the leader3 of 10 ranked benchmarks measured
- 74.50Oct 6, 2026
- SciCode40.86Sep 4, 2026aa_scicode
- Terminal-Bench Hard37.88Oct 7, 2026aa_terminalbench_hard
30.2% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro73.82Oct 7, 2026aa_mmmu_pro
35.7% behind the leader1 of 7 ranked benchmarks measured
- 35.82Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- τ²-Bench Telecom (AA run)86.84Oct 7, 2026aa_tau2
- AA Intelligence24.88Oct 7, 2026aa_intelligence_index
- -7.53Oct 7, 2026
- 38.87Jun 18, 2026
- 52.71Jun 18, 2026
GPT-5 Codex: common questions
Who makes GPT-5 Codex?
GPT-5 Codex is made by OpenAI.
When was GPT-5 Codex released?
GPT-5 Codex was released on Sep 23, 2025, according to Artificial Analysis.
What is GPT-5 Codex good at?
GPT-5 Codex is capable in long context; and behind the leaders in instruction following, factuality, reasoning, coding, multimodal tasks, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does GPT-5 Codex cost?
GPT-5 Codex costs $1.25 per million input tokens and $10.00 per million output tokens, according to Artificial Analysis. At a mix of three input tokens to one output token, it costs more than 78% of the 329 priced models we track.
How many benchmarks has GPT-5 Codex been tested on?
We track 17 results for GPT-5 Codex on 17 benchmarks from 2 sources. The latest was recorded on Oct 7, 2026.
Which API features does GPT-5 Codex support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5 Codex.
About this record
Where GPT-5 Codex's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 28, 2026
Where the results come from
Verification: 17 scores · 0 independently verified · 16 aggregator-attributed · 1 vendor-reported. How these tiers are assigned
From 2 sources on 2 sites. Artificial Analysis supplies 16 of them. Bars are coloured by trust tier.
- artificialanalysis.ai16
- api.llm-stats.com1