GPT-4.1
GPT-4.1 is behind the leaders in long context, agentic tasks, multimodal tasks, and coding. Too few results yet to rate reasoning, safety, math, multilingual tasks, instruction following, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Reasoning, Safety, Math, Multilingual, Instruction Following or Factuality.
Price
$2.00input$8.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
63results on52benchmarks
- 24 independently verified
- 16 aggregator
- 14 vendor-reported
- 9 cross-referenced
From 24 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputs
As listed by OpenRouter
Research
113 papers reference GPT-4.1GPT-4.1 benchmark results
63 results on 52 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
27.0% behind the leader1 of 3 ranked benchmarks measured
- 68.33Oct 8, 2026
40.8% behind the leader1 of 7 ranked benchmarks measured
- 13.84Jun 15, 2026
41.1% behind the leader4 of 6 ranked benchmarks measured
- 74.80Oct 7, 2026
- 72.20Oct 7, 2026
- MMMU-Pro61.21Oct 8, 2026aa_mmmu_pro
- CharXiv (reasoning)56.70Oct 7, 2026CharXiv-R
Show 2 more multimodal resultsHide 2 multimodal results
- 1210.33Aug 25, 2026
- 1214Jun 3, 2026
41.8% behind the leader4 of 10 ranked benchmarks measured
- SciCode38.08Sep 4, 2026aa_scicode
- 54.60Oct 7, 2026
- LiveCodeBench v644.70Jun 15, 2026LiveCodeBench v6 (Aug 24 - May 25)
- Terminal-Bench Hard13.64Oct 8, 2026aa_terminalbench_hard
Show 4 more coding resultsHide 4 coding results
- SWE-bench Multilingual31.50Jun 15, 2026SWE-bench Multilingual (Agentic Coding)
- 39.58Sep 1, 2026
- SWE-bench Verified54.60Jun 15, 2026SWE-bench Verified (Agentic Coding)
- SWE-bench Verified40.80Jun 15, 2026SWE-bench Verified (Agentless Coding)
0 of 6 ranked benchmarks measured
- 27.00May 10, 2026
- 0.42May 10, 2026
- 0.00Oct 8, 2026
Show 4 more reasoning resultsHide 4 reasoning results
- GPQA Diamond66.57Oct 8, 2026gpqa
- GPQA Diamond66.30Oct 7, 2026GPQA
- Humanity's Last Exam4.18Oct 8, 2026aa_hle
- 5.40Oct 7, 2026
0 of 5 ranked benchmarks measured
- 5.96Sep 9, 2026
0 of 3 ranked benchmarks measured
- 39.43Oct 8, 2026
- IFBench42.99Oct 8, 2026aa_ifbench
- 38.30Oct 7, 2026
0 of 4 ranked benchmarks measured
- 31.10Sep 1, 2026
- 5.60May 2, 2026
- AA-Omniscience · Accuracy27.77Oct 8, 2026omniscienceAccuracy
Show 1 more factuality resultHide 1 factuality result
- AA-Omniscience · Non-hallucination6.69Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence13.00Aug 7, 2026Artificial Analysis Intelligence Index
- frontiermath_tier_4_v10.00May 20, 2026frontiermath_tier_4
- 64.79May 19, 2026
- 92.60May 10, 2026
- 97.89May 10, 2026
- 100.00May 10, 2026
Show 26 more resultsHide 26 results
- 27.26Jun 18, 2026
- AA Intelligence12.69Oct 8, 2026aa_intelligence_index
- -39.63Oct 8, 2026
- 52.40May 1, 2026
- 51.60Oct 7, 2026
- 52.40Jun 15, 2026
- 46.40Oct 7, 2026
- 99.25May 10, 2026
- 5.50May 10, 2026
- 21.78Jun 18, 2026
- 91.69May 10, 2026
- 28.90Oct 7, 2026
- 87.40Aug 31, 2026
- 72.70May 26, 2025
- 87.30Oct 7, 2026
- 86.70Jun 15, 2026
- 19.50Jun 15, 2026
- 39.58May 1, 2026
- TAU-bench (airline)49.40Oct 7, 2026TAU-bench Airline
- TAU-bench (retail)68.00Oct 7, 2026TAU-bench Retail
- 30.30Jun 15, 2026
- 8.30Jun 15, 2026
- vectara_answer_rate99.90May 2, 2026Answer Rate
- vectara_avg_summary_length91.70May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency94.40May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)47.08Oct 8, 2026aa_tau2
GPT-4.1: common questions
Who makes GPT-4.1?
GPT-4.1 is made by OpenAI.
When was GPT-4.1 released?
GPT-4.1 was released on Apr 14, 2025, according to Artificial Analysis.
What is GPT-4.1 good at?
GPT-4.1 is behind the leaders in long context, agentic tasks, multimodal tasks, and coding. Too few results yet to rate reasoning, safety, math, multilingual tasks, instruction following, or factuality.
How much does GPT-4.1 cost?
GPT-4.1 costs $2.00 per million input tokens and $8.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 81% of the 330 priced models we track.
How many benchmarks has GPT-4.1 been tested on?
We track 63 results for GPT-4.1 on 52 benchmarks from 24 sources, 24 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-4.1 support?
OpenRouter lists tool calling and structured outputs for GPT-4.1.
About this record
Where GPT-4.1's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 28, 2026
Where the results come from
Verification: 63 scores · 24 independently verified · 16 aggregator-attributed · 9 vendor cross-reference · 14 vendor-reported. How these tiers are assigned
From 24 sources on 14 sites. Artificial Analysis supplies 17 of them; the 24 independently verified results come from 12 sites. Bars are coloured by trust tier.
- artificialanalysis.ai17
- api.llm-stats.com14
- huggingface.co9
- storage.googleapis.com6
- raw.githubusercontent.com4
- epoch.ai3
- arcprize.org2
- swebench.com2
- aider.chat1
- arxiv.org1
- datasets-server.huggingface.co1
- labs.scale.com1
- lmarena.ai1
- simple-bench.com1