Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-4.1

Score basis

GPT-4.1 is behind the leaders in long context, agentic tasks, multimodal tasks, and coding. Too few results yet to rate reasoning, safety, math, multilingual tasks, instruction following, or factuality.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited4 capabilities
  1. Long Context−27.0%1 of 3
  2. Agentic−40.8%1 of 7
  3. Multimodal−41.1%4 of 6
  4. Coding−41.8%4 of 10
Not rated6 capabilities

Too few results yet to rate Reasoning, Safety, Math, Multilingual, Instruction Following or Factuality.

Price

$2.00input$8.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 81% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

63results on52benchmarks

  • 24 independently verified
  • 16 aggregator
  • 14 vendor-reported
  • 9 cross-referenced

From 24 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputs

As listed by OpenRouter

GPT-4.1 benchmark results

63 results on 52 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

LimitedFull long context ranking

27.0% behind the leader1 of 3 ranked benchmarks measured

Agentic

LimitedFull agentic ranking

40.8% behind the leader1 of 7 ranked benchmarks measured

Multimodal

LimitedFull multimodal ranking

41.1% behind the leader4 of 6 ranked benchmarks measured

Show 2 more multimodal resultsHide 2 multimodal results

Coding

LimitedFull coding ranking

41.8% behind the leader4 of 10 ranked benchmarks measured

Show 4 more coding resultsHide 4 coding results

Reasoning

Not enough dataFull reasoning ranking

0 of 6 ranked benchmarks measured

Show 4 more reasoning resultsHide 4 reasoning results

Math

Not enough dataFull math ranking

0 of 5 ranked benchmarks measured

Instruction Following

Not enough dataFull instruction following ranking

0 of 3 ranked benchmarks measured

Factuality

Not enough dataFull factuality ranking

0 of 4 ranked benchmarks measured

Show 1 more factuality resultHide 1 factuality result

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 26 more resultsHide 26 results

GPT-4.1: common questions

Who makes GPT-4.1?

GPT-4.1 is made by OpenAI.

When was GPT-4.1 released?

GPT-4.1 was released on Apr 14, 2025, according to Artificial Analysis.

What is GPT-4.1 good at?

GPT-4.1 is behind the leaders in long context, agentic tasks, multimodal tasks, and coding. Too few results yet to rate reasoning, safety, math, multilingual tasks, instruction following, or factuality.

How much does GPT-4.1 cost?

GPT-4.1 costs $2.00 per million input tokens and $8.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 81% of the 330 priced models we track.

How many benchmarks has GPT-4.1 been tested on?

We track 63 results for GPT-4.1 on 52 benchmarks from 24 sources, 24 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does GPT-4.1 support?

OpenRouter lists tool calling and structured outputs for GPT-4.1.

About this record

Where GPT-4.1's numbers come from, and every name it appears under.

Tracked since
Apr 27, 2026
Newest source mention
Aug 28, 2026

Where the results come from

Verification: 63 scores · 24 independently verified · 16 aggregator-attributed · 9 vendor cross-reference · 14 vendor-reported. How these tiers are assigned

From 24 sources on 14 sites. Artificial Analysis supplies 17 of them; the 24 independently verified results come from 12 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai17
  • api.llm-stats.com14
  • huggingface.co9
  • storage.googleapis.com6
  • raw.githubusercontent.com4
  • epoch.ai3
  • arcprize.org2
  • swebench.com2
  • aider.chat1
  • arxiv.org1
  • datasets-server.huggingface.co1
  • labs.scale.com1
  • lmarena.ai1
  • simple-bench.com1

Also known as

GPT 4.1:batchgpt-4-1gpt-4.1 (2025-04-14)gpt-4.1 (Non-Reasoning)gpt-4.1 (Reasoning)gpt-4.1-2025-04-14gpt4.1