Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

GPT-5.1

Score basis

GPT-5.1 is capable in long context, instruction following, multimodal tasks, factuality, and coding; and behind the leaders in reasoning and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Capable5 capabilities
  1. Long Context−10.7%1 of 3
  2. Instruction Following−19.6%2 of 3
  3. Multimodal−20.8%2 of 6
  4. Factuality−22.0%4 of 4
  5. Coding−23.3%6 of 10
Limited2 capabilities
  1. Reasoning−27.1%5 of 6
  2. Agentic−43.8%4 of 7
Not rated3 capabilities

Too few results yet to rate Safety, Math or Multilingual.

Price

$1.25input$10.00outputper million tokens

From Artificial Analysis · 2 providers tracked · All prices

Costs more than 78% of 329 priced models · 3:1 input-to-output blend, log scale

Evidence

105results on66benchmarks

  • 41 independently verified
  • 34 aggregator
  • 15 vendor-reported
  • 15 cross-referenced

From 29 sources · latest Oct 7, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

GPT-5.1 benchmark results

105 results on 66 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Long Context

CapableFull long context ranking

10.7% behind the leader1 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Instruction Following

CapableFull instruction following ranking

19.6% behind the leader2 of 3 ranked benchmarks measured

Show 1 more instruction following resultHide 1 instruction following result

Multimodal

CapableFull multimodal ranking

20.8% behind the leader2 of 6 ranked benchmarks measured

Show 5 more multimodal resultsHide 5 multimodal results

Factuality

CapableFull factuality ranking

22.0% behind the leader4 of 4 ranked benchmarks measured

Show 2 more factuality resultsHide 2 factuality results

Coding

CapableFull coding ranking

23.3% behind the leader6 of 10 ranked benchmarks measured

Show 7 more coding resultsHide 7 coding results

Reasoning

LimitedFull reasoning ranking

27.1% behind the leader5 of 6 ranked benchmarks measured

Show 9 more reasoning resultsHide 9 reasoning results

Agentic

LimitedFull agentic ranking

43.8% behind the leader4 of 7 ranked benchmarks measured

Show 1 more agentic resultHide 1 agentic result

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 49 more resultsHide 49 results

GPT-5.1: common questions

Who makes GPT-5.1?

GPT-5.1 is made by OpenAI.

When was GPT-5.1 released?

GPT-5.1 was released on Nov 13, 2025, according to Artificial Analysis.

What is GPT-5.1 good at?

GPT-5.1 is capable in long context, instruction following, multimodal tasks, factuality, and coding; and behind the leaders in reasoning and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.

How much does GPT-5.1 cost?

GPT-5.1 costs $1.25 per million input tokens and $10.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 78% of the 329 priced models we track.

How many benchmarks has GPT-5.1 been tested on?

We track 105 results for GPT-5.1 on 66 benchmarks from 29 sources, 41 of them independently verified. The latest was recorded on Oct 7, 2026.

Which API features does GPT-5.1 support?

OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.1.

About this record

Where GPT-5.1's numbers come from, and every name it appears under.

Tracked since
Jun 20, 2026
Newest source mention
Sep 7, 2026

Where the results come from

Verification: 105 scores · 41 independently verified · 34 aggregator-attributed · 15 vendor cross-reference · 15 vendor-reported. How these tiers are assigned

From 29 sources on 15 sites. Artificial Analysis supplies 36 of them; the 41 independently verified results come from 11 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai36
  • huggingface.co9
  • arcprize.org8
  • deploymentsafety.openai.com8
  • api.llm-stats.com7
  • storage.googleapis.com6
  • www-cdn.anthropic.com6
  • datasets-server.huggingface.co5
  • raw.githubusercontent.com5
  • matharena.ai4
  • epoch.ai3
  • lmarena.ai3
  • labs.scale.com2
  • swebench.com2
  • simple-bench.com1

Also known as

How our sources name GPT-5.1 at each reasoning setting.

SettingShort formAPI id
low—gpt-5.1 (thinking, low) gpt-5.1-low-2025-11-13
mediumGPT 5.1 (2025-11-13) (medium)gpt-5.1 (2025-11-13) (medium reasoning) gpt-5.1 (medium) gpt-5.1 (thinking, medium) gpt-5.1-medium
high—gpt-5.1 (high) gpt-5.1 (thinking, high) gpt-5.1-high gpt-5.1-high-2025-11-13
Also listed asgpt-5-1gpt-5-1 (Non-Reasoning)gpt-5-1 (reasoning)gpt-5-1-non-reasoninggpt-5.1 (non-reasoning)gpt-5.1 (thinking, none)gpt-5.1-2025-11-13gpt-5.1-2025-11-13-thinkinggpt-5.1-20251113gpt-5.1-thinkinggpt-5.1:batchgpt-5.1 (2025-11-13)