GPT-5.4
GPT-5.4 is strong in long context; capable in reasoning, coding, instruction following, multimodal tasks, agentic tasks, and math; and behind the leaders in factuality. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$2.50input$15.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
186results on109benchmarks
- 41 independently verified
- 53 aggregator
- 26 vendor-reported
- 66 cross-referenced
From 28 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
206 papers reference GPT-5.4GPT-5.4 benchmark results
186 results on 109 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
8.4% behind the leader1 of 3 ranked benchmarks measured
- 82.00Oct 8, 2026
11.0% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond92.02Oct 8, 2026gpqa
- LiveBench · Reasoning88.12Oct 8, 2026livebench_reasoning@2026-06-25
- ARC-AGI-273.30Oct 7, 2026ARC-AGI v2
- 23.43Oct 8, 2026
- Humanity's Last Exam43.74Oct 8, 2026aa_hle
Show 18 more reasoning resultsHide 18 reasoning results
- 73.95Sep 21, 2026
- 67.50May 10, 2026
- 55.42May 10, 2026
- 29.17May 10, 2026
- 73.30Jun 21, 2026
- 0.21May 10, 2026
- 0.57Oct 8, 2026
- 7.43Oct 8, 2026
- GPQA Diamond87.07Oct 8, 2026gpqa
- GPQA Diamond74.85Oct 8, 2026gpqa
- GPQA Diamond92.80Oct 7, 2026GPQA
- 92.00Oct 8, 2026
- GPQA Diamond93.00Jul 4, 2026GPQA Diamond (Pass@1)
- 92.80Jun 21, 2026
- Humanity's Last Exam30.82Oct 8, 2026aa_hle
- Humanity's Last Exam11.26Oct 8, 2026aa_hle
- 39.80Oct 7, 2026
- 39.80Oct 8, 2026
20.7% behind the leader8 of 10 ranked benchmarks measured
- Terminal-Bench Hard57.58Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.178.28Oct 8, 2026terminalbenchV21
- LiveBench · Coding77.54Oct 8, 2026livebench_coding@2026-06-25
- SciCode56.60Sep 4, 2026aa_scicode
- 71.70Jun 15, 2026
- LiveBench · Agentic Coding53.84Oct 8, 2026livebench_agentic_coding@2026-06-25
- 1465.23May 22, 2026
- 57.70Oct 7, 2026
Show 10 more coding resultsHide 10 coding results
- 1444.37Sep 21, 2026
- 1399.22May 22, 2026
- SciCode47.11Sep 4, 2026aa_scicode
- SciCode50.35Sep 4, 2026aa_scicode
- 56.60May 3, 2026
- 59.10Oct 8, 2026
- 57.70Oct 8, 2026
- 46.00Jun 15, 2026
- Terminal-Bench Hard43.18Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard37.88Oct 8, 2026aa_terminalbench_hard
22.4% behind the leader2 of 3 ranked benchmarks measured
- IFBench73.95Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following70.22Oct 8, 2026livebench_instruction_following@2026-06-25
22.9% behind the leader2 of 6 ranked benchmarks measured
- MMMU-Pro78.44Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)82.80May 3, 2026CharXiv (RQ)
Show 8 more multimodal resultsHide 8 multimodal results
23.8% behind the leader5 of 7 ranked benchmarks measured
- 82.70Oct 7, 2026
- 75.00Oct 7, 2026
- τ-Bench V3 · Banking39.59Oct 8, 2026tauBanking
- 67.20Oct 7, 2026
- 37.42Oct 8, 2026
Show 11 more agentic resultsHide 11 agentic results
- 33.26Oct 8, 2026
- AA ApexAgents33.30May 3, 2026APEX-Agents
- 34.50Jul 10, 2026
- 18.90Jul 10, 2026
- BrowseComp82.70Jul 4, 2026BrowseComp (Pass@1)
- 50.15Jun 15, 2026
- 42.12Jun 15, 2026
- 70.60Oct 8, 2026
- MCP Atlas67.20Oct 8, 2026MCP-Atlas (Public Set)
- 68.10Jun 21, 2026
- OSWorld-Verified75.00Jun 21, 2026OSWorld
24.8% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics94.15Oct 8, 2026livebench_math@2026-06-25
- 78.60Sep 21, 2026
- 49.00Sep 21, 2026
Show 10 more math resultsHide 10 math results
- 99.17Sep 2, 2026
- 99.17May 2, 2026
- 98.70Oct 8, 2026
- 99.20May 4, 2026
- 97.73Sep 2, 2026
- 97.73May 10, 2026
- HMMT Feb 202691.80Oct 8, 2026HMMT Feb. 2026
- HMMT Feb 202697.70Jul 4, 2026HMMT 2026 Feb (Pass@1)
- 91.40Oct 8, 2026
- 95.24Sep 2, 2026
28.2% behind the leader4 of 4 ranked benchmarks measured
- 7.00May 2, 2026
- AA-Omniscience · Accuracy50.85Oct 8, 2026omniscienceAccuracy
- 45.10May 20, 2026
- AA-Omniscience · Non-hallucination8.31Oct 8, 2026omniscienceNonHallucination
Show 5 more factuality resultsHide 5 factuality results
- AA-Omniscience · Accuracy37.38Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy47.85Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination17.42Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination15.28Oct 8, 2026omniscienceNonHallucination
- SimpleQA Verified45.30Jun 27, 2026SimpleQA-Verified (Pass@1)
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_language82.63Oct 8, 2026livebench_language@2026-06-25
- livebench_data_analysis79.31Oct 8, 2026livebench_data_analysis@2026-06-25
- 9.67Oct 8, 2026
- 1475Oct 5, 2026
- 93.67Sep 21, 2026
- 92.47Sep 2, 2026
Show 84 more resultsHide 84 results
- 44.17Sep 4, 2026
- 58.22Jun 18, 2026
- 39.14Jun 18, 2026
- AA Intelligence38.98Oct 8, 2026aa_intelligence_index
- AA Intelligence27.58Oct 8, 2026aa_intelligence_index
- AA Intelligence18.16Oct 8, 2026aa_intelligence_index
- 5.78Oct 8, 2026
- 4.78Oct 8, 2026
- -15.67Oct 8, 2026
- 60.00Jun 24, 2026
- 70.10Jun 24, 2026
- 68.58Jun 24, 2026
- 58.25Jun 24, 2026
- 37.26Jun 24, 2026
- 66.29Jun 24, 2026
- 53.69Jun 24, 2026
- 51.80Jun 24, 2026
- 54.10Jul 4, 2026
- 78.10Jul 4, 2026
- 92.67May 10, 2026
- 86.17May 10, 2026
- 68.17May 10, 2026
- 93.70Jun 21, 2026
- Artificial Analysis Coding Index71.05Sep 9, 2026aa_coding_index
- 45.57Jun 18, 2026
- 40.95Jun 18, 2026
- baby_vision_with_python80.20May 3, 2026BabyVision (w/ python)
- 49.70May 3, 2026
- 82.70May 3, 2026
- browsecomp_with_context_manager82.70Oct 8, 2026BrowseComp (w/ Context Manage)
- charxiv_rq_with_python90.00May 3, 2026CharXiv (RQ) (w/ python)
- Chinese SimpleQA (C-SimpleQA)76.80Jul 4, 2026Chinese-SimpleQA (Pass@1)
- 78.40May 3, 2026
- 3168.00Jul 4, 2026
- 66.30Oct 8, 2026
- 63.70May 3, 2026
- DeepSearchQA (F1)78.60May 3, 2026DeepSearchQA (f1-score)
- 52.00Oct 7, 2026
- 12.82Jul 7, 2026
- 56.00Oct 7, 2026
- 57.20Jun 21, 2026
- FrontierMath (overall)47.60Oct 7, 2026FrontierMath
- frontiermath_tier_4_v127.10Aug 29, 2026frontiermath_tier_4
- 1674.00Jul 4, 2026
- 3.50Aug 11, 2026
- 96.30Jul 9, 2026
- 29.10Jul 9, 2026
- 54.00Jul 9, 2026
- 48.10Jul 9, 2026
- HLE (with tools)52.10Oct 8, 2026HLE (w/ Tools)
- HLE (with tools)52.00Jul 4, 2026HLE w/ tools (Pass@1)
- 95.80Oct 8, 2026
- Image input eval - extremism (not_unsafe)0.99Jul 9, 2026
- Image input eval - harms-erotic (not_unsafe)0.99Jul 9, 2026
- Image input eval - hate (not_unsafe)0.99Jul 9, 2026
- Image input eval - self-harm (not_unsafe)1.00Jul 9, 2026
- 0.40Oct 7, 2026
- 80.28Oct 7, 2026
- 92.00May 3, 2026
- mathvision_with_python96.10May 3, 2026MathVision (w/ python)
- 67.20Jul 4, 2026
- 62.50May 3, 2026
- MMLU-Pro87.50Jul 4, 2026MMLU-Pro (EM)
- mmmu_pro_with_python82.10May 3, 2026MMMU-Pro (w/ python)
- 41.30Oct 8, 2026
- 68.10Jun 21, 2026
- 51.10Jun 21, 2026
- 89.10Oct 7, 2026
- 81.30Jun 15, 2026
- 75.10Oct 7, 2026
- Terminal-Bench 2.075.10Oct 8, 2026Terminal-Bench 2.0 (Best self-reported)
- Terminal-Bench 2.065.40May 3, 2026Terminal-Bench 2.0 (Terminus-2)
- 54.60Oct 8, 2026
- 54.60Oct 7, 2026
- Toolathlon54.60Jul 4, 2026Toolathlon (Pass@1)
- 98.40May 3, 2026
- vectara_answer_rate99.90May 2, 2026Answer Rate
- vectara_avg_summary_length81.70May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency93.00May 2, 2026Factual Consistency Rate
- 6144.18Oct 8, 2026
- τ²-Bench Telecom (AA run)87.13Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)74.56Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)35.96Oct 8, 2026aa_tau2
- 72.90Oct 8, 2026
GPT-5.4: common questions
Who makes GPT-5.4?
GPT-5.4 is made by OpenAI.
When was GPT-5.4 released?
GPT-5.4 was released on Mar 5, 2026, according to Artificial Analysis.
What is GPT-5.4 good at?
GPT-5.4 is strong in long context; capable in reasoning, coding, instruction following, multimodal tasks, agentic tasks, and math; and behind the leaders in factuality. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5.4 cost?
GPT-5.4 costs $2.50 per million input tokens and $15.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 88% of the 330 priced models we track.
How many benchmarks has GPT-5.4 been tested on?
We track 186 results for GPT-5.4 on 109 benchmarks from 28 sources, 41 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.4 support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.4.
About this record
Where GPT-5.4's numbers come from, and every name it appears under.
- Tracked since
- Apr 25, 2026
- Newest source mention
- Aug 24, 2026
Where the results come from
Verification: 186 scores · 41 independently verified · 53 aggregator-attributed · 66 vendor cross-reference · 26 vendor-reported. How these tiers are assigned
From 28 sources on 13 sites. Hugging Face supplies 58 of them; the 41 independently verified results come from 8 sites. Bars are coloured by trust tier.
- huggingface.co58
- artificialanalysis.ai53
- api.llm-stats.com16
- deploymentsafety.openai.com10
- arcprize.org9
- www-cdn.anthropic.com8
- livebench.ai7
- matharena.ai6
- datasets-server.huggingface.co5
- epoch.ai4
- raw.githubusercontent.com4
- labs.scale.com3
- lmarena.ai3
Also known as
How our sources name GPT-5.4 at each reasoning setting.
| Setting | Short form | Long form | API id |
|---|---|---|---|
| low | — | — | gpt-5-4-low gpt-5.4 (low) |
| medium | — | — | gpt-5.4 (medium) gpt-5.4-medium (codex-harness) |
| high | — | gpt-5.4-2026-03-05 (reasoning effort = high) | gpt-5.4 (high) gpt-5.4-high gpt-5.4-high (codex-harness) |
| xhigh | gpt-5.4 xhigh | — | gpt-5.4 (xhigh) gpt-5.4 (xhigh)* gpt-5.4-2026-03-05 (xhigh thinking) |