Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.4

Score basis

GPT-5.4 is strong in long context; capable in reasoning, coding, instruction following, multimodal tasks, agentic tasks, and math; and behind the leaders in factuality. Too few results yet to rate safety or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Strong1 capability
  1. Long Context−8.4%1 of 3
Capable6 capabilities
  1. Reasoning−11.0%5 of 6
  2. Coding−20.7%8 of 10
  3. Instruction Following−22.4%2 of 3
  4. Multimodal−22.9%2 of 6
  5. Agentic−23.8%5 of 7
  6. Math−24.8%3 of 5
Limited1 capability
  1. Factuality−28.2%4 of 4
Not rated2 capabilities

Too few results yet to rate Safety or Multilingual.

Price

$2.50input$15.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 88% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

186results on109benchmarks

  • 41 independently verified
  • 53 aggregator
  • 26 vendor-reported
  • 66 cross-referenced

From 28 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

GPT-5.4 benchmark results

186 results on 109 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

StrongFull long context ranking

8.4% behind the leader1 of 3 ranked benchmarks measured

Show 2 more long context resultsHide 2 long context results

Reasoning

CapableFull reasoning ranking

11.0% behind the leader5 of 6 ranked benchmarks measured

Show 18 more reasoning resultsHide 18 reasoning results

Coding

CapableFull coding ranking

20.7% behind the leader8 of 10 ranked benchmarks measured

Show 10 more coding resultsHide 10 coding results

Instruction Following

CapableFull instruction following ranking

22.4% behind the leader2 of 3 ranked benchmarks measured

Show 2 more instruction following resultsHide 2 instruction following results

Multimodal

CapableFull multimodal ranking

22.9% behind the leader2 of 6 ranked benchmarks measured

Show 8 more multimodal resultsHide 8 multimodal results

Agentic

CapableFull agentic ranking

23.8% behind the leader5 of 7 ranked benchmarks measured

Show 11 more agentic resultsHide 11 agentic results

24.8% behind the leader3 of 5 ranked benchmarks measured

Show 10 more math resultsHide 10 math results

Factuality

LimitedFull factuality ranking

28.2% behind the leader4 of 4 ranked benchmarks measured

Show 5 more factuality resultsHide 5 factuality results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 84 more resultsHide 84 results

GPT-5.4: common questions

Who makes GPT-5.4?

GPT-5.4 is made by OpenAI.

When was GPT-5.4 released?

GPT-5.4 was released on Mar 5, 2026, according to Artificial Analysis.

What is GPT-5.4 good at?

GPT-5.4 is strong in long context; capable in reasoning, coding, instruction following, multimodal tasks, agentic tasks, and math; and behind the leaders in factuality. Too few results yet to rate safety or multilingual tasks.

How much does GPT-5.4 cost?

GPT-5.4 costs $2.50 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.

How many benchmarks has GPT-5.4 been tested on?

We track 186 results for GPT-5.4 on 109 benchmarks from 28 sources, 41 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does GPT-5.4 support?

OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.4.

About this record

Where GPT-5.4's numbers come from, and every name it appears under.

Tracked since
Apr 25, 2026
Newest source mention
Aug 24, 2026

Where the results come from

Verification: 186 scores · 41 independently verified · 53 aggregator-attributed · 66 vendor cross-reference · 26 vendor-reported. How these tiers are assigned

From 28 sources on 13 sites. Hugging Face supplies 58 of them; the 41 independently verified results come from 8 sites. Bars are coloured by trust tier.

  • huggingface.co58
  • artificialanalysis.ai53
  • api.llm-stats.com16
  • deploymentsafety.openai.com10
  • arcprize.org9
  • www-cdn.anthropic.com8
  • livebench.ai7
  • matharena.ai6
  • datasets-server.huggingface.co5
  • epoch.ai4
  • raw.githubusercontent.com4
  • labs.scale.com3
  • lmarena.ai3

Also known as

How our sources name GPT-5.4 at each reasoning setting.

SettingShort formLong formAPI id
low——gpt-5-4-low gpt-5.4 (low)
medium——gpt-5.4 (medium) gpt-5.4-medium (codex-harness)
high—gpt-5.4-2026-03-05 (reasoning effort = high)gpt-5.4 (high) gpt-5.4-high gpt-5.4-high (codex-harness)
xhighgpt-5.4 xhigh—gpt-5.4 (xhigh) gpt-5.4 (xhigh)* gpt-5.4-2026-03-05 (xhigh thinking)
Also listed asgpt-5-4gpt-5-4-non-reasoninggpt-5.4 (non-reasoning)gpt-5.4 (reasoning)gpt-5.4 (w/o explore)gpt-5.4 thinkinggpt-5.4 w/o exploregpt-5.4-2026-03-05gpt-5.4:batchgpt5.4