GPT-4.1 mini
GPT-4.1 mini is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Coding, Agentic, Safety, Long Context, Math, Multilingual, Instruction Following or Factuality.
Price
$0.40input$1.60outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
48results on42benchmarks
- 16 independently verified
- 18 aggregator
- 14 vendor-reported
From 17 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Research
41 papers reference GPT-4.1 miniGPT-4.1 mini benchmark results
48 results on 42 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
42.9% behind the leader4 of 6 ranked benchmarks measured
- 72.70Oct 7, 2026
- 73.10Oct 7, 2026
- MMMU-Pro58.73Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)56.80Oct 7, 2026CharXiv-R
Show 2 more multimodal resultsHide 2 multimodal results
- 1181.26Aug 25, 2026
- 1203May 6, 2026
0 of 6 ranked benchmarks measured
- 0.00May 10, 2026
- 0.00Oct 8, 2026
- GPQA Diamond66.36Oct 8, 2026gpqa
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond65.00Oct 7, 2026GPQA
- Humanity's Last Exam5.02Oct 8, 2026aa_hle
- 3.70Oct 7, 2026
0 of 10 ranked benchmarks measured
- 23.94Sep 1, 2026
- Terminal-Bench 2.110.11Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard7.58Oct 8, 2026aa_terminalbench_hard
Show 2 more coding resultsHide 2 coding results
- SciCode40.39Sep 4, 2026aa_scicode
- 23.60Oct 7, 2026
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- τ-Bench V3 · Banking5.36Oct 8, 2026tauBanking
0 of 3 ranked benchmarks measured
- 44.00Oct 8, 2026
0 of 5 ranked benchmarks measured
- 6.67Sep 9, 2026
0 of 3 ranked benchmarks measured
- IFBench38.30Oct 8, 2026aa_ifbench
- 35.80Oct 7, 2026
0 of 4 ranked benchmarks measured
- 12.70Sep 1, 2026
- AA-Omniscience · Accuracy20.32Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination7.28Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence10.00Aug 7, 2026Artificial Analysis Intelligence Index
- 32.40Jun 19, 2026
- 23.94Jun 19, 2026
- 60.44May 19, 2026
- 92.10May 10, 2026
- 97.44May 10, 2026
Show 16 more resultsHide 16 results
- 1.79Sep 4, 2026
- AA Intelligence10.16Oct 8, 2026aa_intelligence_index
- -53.57Oct 8, 2026
- 34.70Oct 7, 2026
- 40.20Oct 7, 2026
- 99.29May 10, 2026
- 3.50May 10, 2026
- Artificial Analysis Coding Index20.21Sep 9, 2026aa_coding_index
- 85.63May 10, 2026
- 35.00Oct 7, 2026
- 84.10Aug 31, 2026
- 78.50Oct 7, 2026
- 100.00May 10, 2026
- TAU-bench (airline)36.00Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)55.80Oct 7, 2026TAU-bench Retail
- τ²-Bench Telecom (AA run)52.92Oct 8, 2026aa_tau2
GPT-4.1 mini: common questions
Who makes GPT-4.1 mini?
GPT-4.1 mini is made by OpenAI.
When was GPT-4.1 mini released?
GPT-4.1 mini was released on Apr 14, 2025, according to Artificial Analysis.
What is GPT-4.1 mini good at?
GPT-4.1 mini is behind the leaders in multimodal tasks. Too few results yet to rate reasoning, coding, agentic tasks, safety, long context, math, multilingual tasks, instruction following, or factuality.
How much does GPT-4.1 mini cost?
GPT-4.1 mini costs $0.40 per million input tokens and $1.60 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 50% of the 331 priced models we track.
How many benchmarks has GPT-4.1 mini been tested on?
We track 48 results for GPT-4.1 mini on 42 benchmarks from 17 sources, 16 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-4.1 mini support?
OpenRouter lists tool calling and structured outputs for GPT-4.1 mini.
About this record
Where GPT-4.1 mini's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 25, 2026
Where the results come from
Verification: 48 scores · 16 independently verified · 18 aggregator-attributed · 14 vendor-reported. How these tiers are assigned
From 17 sources on 9 sites. Artificial Analysis supplies 19 of them; the 16 independently verified results come from 8 sites. Bars are coloured by trust tier.
- artificialanalysis.ai19
- api.llm-stats.com14
- storage.googleapis.com6
- arcprize.org2
- epoch.ai2
- swebench.com2
- aider.chat1
- datasets-server.huggingface.co1
- lmarena.ai1