GPT-5.4 mini
GPT-5.4 mini is capable in long context; and behind the leaders in reasoning, multimodal tasks, agentic tasks, coding, instruction following, and math. Too few results yet to rate safety, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Multilingual or Factuality.
Price
$0.75input$4.50outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
95results on49benchmarks
- 30 independently verified
- 54 aggregator
- 11 vendor-reported
From 13 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
54 papers reference GPT-5.4 miniGPT-5.4 mini benchmark results
95 results on 49 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
16.4% behind the leader1 of 3 ranked benchmarks measured
- 77.00Oct 8, 2026
27.7% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond87.47Oct 8, 2026gpqa
- LiveBench · Reasoning71.32Oct 8, 2026livebench_reasoning@2026-06-25
- Humanity's Last Exam28.13Oct 8, 2026aa_hle
- 10.00Oct 8, 2026
- 18.90May 10, 2026
Show 11 more reasoning resultsHide 11 reasoning results
- 1.11Sep 21, 2026
- 13.19Sep 21, 2026
- 4.44May 15, 2026
- 0.00Oct 8, 2026
- 2.86Oct 8, 2026
- GPQA Diamond60.61Oct 8, 2026gpqa
- GPQA Diamond82.32Oct 8, 2026gpqa
- GPQA Diamond88.00Oct 7, 2026GPQA
- Humanity's Last Exam5.89Oct 8, 2026aa_hle
- Humanity's Last Exam18.58Oct 8, 2026aa_hle
- 28.20Oct 7, 2026
30.6% behind the leader1 of 6 ranked benchmarks measured
- MMMU-Pro73.29Oct 8, 2026aa_mmmu_pro
Show 5 more multimodal resultsHide 5 multimodal results
35.3% behind the leader6 of 7 ranked benchmarks measured
- 72.10Oct 7, 2026
- 57.70Oct 7, 2026
- τ-Bench V3 · Banking25.57Oct 8, 2026tauBanking
- 25.80Oct 8, 2026
- Terminal-Bench 4.02.02Oct 8, 2026
Show 6 more agentic resultsHide 6 agentic results
- 28.17Oct 8, 2026
- 18.64Oct 8, 2026
- 35.22Oct 8, 2026
- 4.65Oct 8, 2026
- 40.76Jun 15, 2026
- 56.70Oct 8, 2026
35.6% behind the leader7 of 10 ranked benchmarks measured
- Terminal-Bench Hard52.27Oct 8, 2026aa_terminalbench_hard
- LiveBench · Coding71.62Oct 8, 2026livebench_coding@2026-06-25
- SciCode52.08Oct 8, 2026aa_scicode
- Terminal-Bench 2.159.18Oct 8, 2026terminalbenchV21
- 54.40Oct 7, 2026
- 1397.25Sep 21, 2026
- LiveBench · Agentic Coding41.67Oct 8, 2026livebench_agentic_coding@2026-06-25
Show 4 more coding resultsHide 4 coding results
- SciCode39.58Sep 4, 2026aa_scicode
- SciCode44.21Sep 4, 2026aa_scicode
- Terminal-Bench Hard18.18Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard34.09Oct 8, 2026aa_terminalbench_hard
36.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench73.27Oct 8, 2026aa_ifbench
- LiveBench · Instruction Following59.80Oct 8, 2026livebench_instruction_following@2026-06-25
52.1% behind the leader3 of 5 ranked benchmarks measured
- LiveBench · Mathematics78.46Oct 8, 2026livebench_math@2026-06-25
- 51.23Sep 21, 2026
- 9.76Sep 21, 2026
Show 2 more math resultsHide 2 math results
- 17.19Sep 22, 2026
- 24.56Sep 21, 2026
0 of 4 ranked benchmarks measured
- 29.40May 20, 2026
- 5.50May 2, 2026
- AA-Omniscience · Accuracy37.47Oct 8, 2026omniscienceAccuracy
Show 5 more factuality resultsHide 5 factuality results
- AA-Omniscience · Accuracy25.65Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy36.47Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination4.42Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination9.83Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination10.15Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- livebench_data_analysis70.79Oct 8, 2026livebench_data_analysis@2026-06-25
- livebench_language70.95Oct 8, 2026livebench_language@2026-06-25
- 13.00Sep 21, 2026
- 58.00Sep 21, 2026
- frontiermath_tier_4_v12.08Aug 29, 2026frontiermath_tier_4
- 1449Jul 23, 2026
Show 25 more resultsHide 25 results
- 19.65Sep 9, 2026
- 25.01Jun 18, 2026
- 40.27Jun 18, 2026
- AA Intelligence11.14Oct 8, 2026aa_intelligence_index
- AA Intelligence24.07Oct 8, 2026aa_intelligence_index
- AA Intelligence19.75Oct 8, 2026aa_intelligence_index
- -45.42Oct 8, 2026
- -18.92Oct 8, 2026
- -20.62Oct 8, 2026
- 40.83May 15, 2026
- 63.67May 10, 2026
- Artificial Analysis Coding Index56.08Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index25.32Jun 18, 2026aa_coding_index
- 37.46Jun 18, 2026
- 45.36Oct 7, 2026
- 0.00Oct 7, 2026
- 87.37Oct 7, 2026
- 60.00Oct 7, 2026
- 42.90Oct 7, 2026
- vectara_answer_rate100.00May 2, 2026Answer Rate
- vectara_avg_summary_length54.70May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency94.50May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)23.39Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)83.33Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)36.55Oct 8, 2026aa_tau2
GPT-5.4 mini: common questions
Who makes GPT-5.4 mini?
GPT-5.4 mini is made by OpenAI.
When was GPT-5.4 mini released?
GPT-5.4 mini was released on Mar 17, 2026, according to Artificial Analysis.
What is GPT-5.4 mini good at?
GPT-5.4 mini is capable in long context; and behind the leaders in reasoning, multimodal tasks, agentic tasks, coding, instruction following, and math. Too few results yet to rate safety, multilingual tasks, or factuality.
How much does GPT-5.4 mini cost?
GPT-5.4 mini costs $0.75 per million input tokens and $4.50 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 68% of the 330 priced models we track.
How many benchmarks has GPT-5.4 mini been tested on?
We track 95 results for GPT-5.4 mini on 49 benchmarks from 13 sources, 30 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5.4 mini support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.4 mini.
About this record
Where GPT-5.4 mini's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 9, 2026
Where the results come from
Verification: 95 scores · 30 independently verified · 54 aggregator-attributed · 11 vendor-reported. How these tiers are assigned
From 13 sources on 9 sites. Artificial Analysis supplies 54 of them; the 30 independently verified results come from 7 sites. Bars are coloured by trust tier.
- artificialanalysis.ai54
- api.llm-stats.com11
- arcprize.org8
- livebench.ai7
- epoch.ai6
- raw.githubusercontent.com4
- datasets-server.huggingface.co2
- lmarena.ai2
- labs.scale.com1
Also known as
How our sources name GPT-5.4 mini at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | gpt-5.4 mini (low) | — |
| medium | gpt-5.4 mini (medium) | gpt-5-4-mini-medium |
| high | gpt-5.4 mini (high) | gpt-5.4-mini-high |
| xhigh | gpt-5.4 mini (xhigh) | — |