GPT-5.1 Codex
GPT-5.1 Codex is behind the leaders in long context, factuality, reasoning, instruction following, multimodal tasks, coding, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$1.25input$10.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
24results on20benchmarks
- 4 independently verified
- 16 aggregator
- 2 vendor-reported
- 2 cross-referenced
From 5 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
GPT-5.1 Codex benchmark results
24 results on 20 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
26.2% behind the leader1 of 3 ranked benchmarks measured
- 69.33Oct 8, 2026
26.6% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy39.90Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination22.80Oct 8, 2026omniscienceNonHallucination
27.6% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond85.96Oct 8, 2026gpqa
- Humanity's Last Exam25.72Oct 8, 2026aa_hle
- 5.71Oct 8, 2026
28.8% behind the leader1 of 3 ranked benchmarks measured
- IFBench70.00Oct 8, 2026aa_ifbench
31.6% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro72.49Oct 8, 2026aa_mmmu_pro
33.2% behind the leader4 of 10 ranked benchmarks measured
- 73.70Oct 7, 2026
- SciCode40.16Sep 4, 2026aa_scicode
- Terminal-Bench Hard34.85Oct 8, 2026aa_terminalbench_hard
- 1335.84May 22, 2026
Show 3 more coding resultsHide 3 coding results
- 1329.73May 30, 2026
- 66.00Sep 25, 2026
- 73.70Jun 15, 2026
36.2% behind the leader1 of 7 ranked benchmarks measured
- 34.55Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 66.00Aug 29, 2026
- τ²-Bench Telecom (AA run)83.04Oct 8, 2026aa_tau2
- AA Intelligence23.70Oct 8, 2026aa_intelligence_index
- -6.50Oct 8, 2026
- 36.62Jun 18, 2026
- 50.68Jun 18, 2026
Show 2 more resultsHide 2 results
- 52.80Oct 7, 2026
- Terminal-Bench 2.052.80Jun 15, 2026Terminal Bench 2
GPT-5.1 Codex: common questions
Who makes GPT-5.1 Codex?
GPT-5.1 Codex is made by OpenAI.
When was GPT-5.1 Codex released?
GPT-5.1 Codex was released on Nov 13, 2025, according to Artificial Analysis.
What is GPT-5.1 Codex good at?
GPT-5.1 Codex is behind the leaders in long context, factuality, reasoning, instruction following, multimodal tasks, coding, and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does GPT-5.1 Codex cost?
GPT-5.1 Codex costs $1.25 per million input tokens and $10.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 78% of the 330 priced models we track.
How many benchmarks has GPT-5.1 Codex been tested on?
We track 24 results for GPT-5.1 Codex on 20 benchmarks from 5 sources, 4 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.1 Codex support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.1 Codex.
About this record
Where GPT-5.1 Codex's numbers come from, and every name it appears under.
- Tracked since
- May 2, 2026
- Newest source mention
- Aug 10, 2026
Where the results come from
Verification: 24 scores · 4 independently verified · 16 aggregator-attributed · 2 vendor cross-reference · 2 vendor-reported. How these tiers are assigned
From 5 sources on 5 sites. Artificial Analysis supplies 16 of them; the 4 independently verified results come from 2 sites. Bars are coloured by trust tier.
- artificialanalysis.ai16
- api.llm-stats.com2
- datasets-server.huggingface.co2
- huggingface.co2
- swebench.com2
Also known as
How our sources name GPT-5.1 Codex at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| medium | gpt 5.1 codex (medium) | gpt-5.1-codex (medium reasoning) |
| high | gpt-5.1 codex (high) gpt 5.1 codex high | — |