GPT-5.1
GPT-5.1 is capable in long context, instruction following, multimodal tasks, factuality, and coding; and behind the leaders in reasoning and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety, Math or Multilingual.
Price
$1.25input$10.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
105results on66benchmarks
- 41 independently verified
- 34 aggregator
- 15 vendor-reported
- 15 cross-referenced
From 29 sources · latest Oct 7, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
35 papers reference GPT-5.1GPT-5.1 benchmark results
105 results on 66 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
10.7% behind the leader1 of 3 ranked benchmarks measured
- 80.00Oct 7, 2026
Show 1 more long context resultHide 1 long context result
- 45.00Oct 7, 2026
19.6% behind the leader2 of 3 ranked benchmarks measured
- IFBench72.86Oct 7, 2026aa_ifbench
- 63.41Oct 7, 2026
Show 1 more instruction following resultHide 1 instruction following result
- IFBench43.20Oct 7, 2026aa_ifbench
20.8% behind the leader2 of 6 ranked benchmarks measured
- 85.40Oct 6, 2026
- MMMU-Pro75.49Oct 7, 2026aa_mmmu_pro
Show 5 more multimodal resultsHide 5 multimodal results
- 1234.90Sep 22, 2026
- 1249.80Sep 21, 2026
- 1238Jun 17, 2026
- 1250May 25, 2026
- MMMU-Pro62.37Oct 7, 2026aa_mmmu_pro
22.0% behind the leader4 of 4 ranked benchmarks measured
- 12.10Aug 29, 2026
- 48.00May 20, 2026
- AA-Omniscience · Accuracy37.72Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination48.11Oct 7, 2026omniscienceNonHallucination
Show 2 more factuality resultsHide 2 factuality results
- AA-Omniscience · Accuracy29.47Oct 7, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination9.38Oct 7, 2026omniscienceNonHallucination
23.3% behind the leader6 of 10 ranked benchmarks measured
- 87.00May 18, 2026
- 76.30Oct 6, 2026
- Terminal-Bench Hard45.45Oct 7, 2026aa_terminalbench_hard
- SciCode43.29Sep 4, 2026aa_scicode
- Terminal-Bench 2.152.43Oct 7, 2026terminalbenchV21
- 1394.79Sep 21, 2026
Show 7 more coding resultsHide 7 coding results
- 1340.99Sep 22, 2026
- 1421.20May 22, 2026
- SciCode36.46Sep 4, 2026aa_scicode
- 66.00Sep 25, 2026
- 76.30Jul 29, 2026
- Terminal-Bench Hard22.73Oct 7, 2026aa_terminalbench_hard
- 43.00May 18, 2026
27.1% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond87.27Oct 7, 2026gpqa
- 53.20May 10, 2026
- Humanity's Last Exam28.50Oct 7, 2026aa_hle
- 17.64May 10, 2026
- 4.86Oct 7, 2026
Show 9 more reasoning resultsHide 9 reasoning results
- 6.53Sep 21, 2026
- 1.94May 10, 2026
- 0.42May 10, 2026
- 0.00Oct 7, 2026
- GPQA Diamond64.34Oct 7, 2026gpqa
- GPQA Diamond88.10Oct 6, 2026GPQA
- 88.10Jul 29, 2026
- Humanity's Last Exam5.28Oct 7, 2026aa_hle
- 25.70May 18, 2026
43.8% behind the leader4 of 7 ranked benchmarks measured
- 50.10Oct 7, 2026
- 50.80May 18, 2026
- τ-Bench V3 · Banking15.88Oct 7, 2026tauBanking
- 16.51Oct 7, 2026
Show 1 more agentic resultHide 1 agentic result
- 25.00Jun 15, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- AA Intelligence25.00Sep 21, 2026Artificial Analysis Intelligence Index
- 57.67Sep 21, 2026
- 98.22Sep 7, 2026
- 99.75Sep 7, 2026
- 86.19Sep 7, 2026
- 97.56Sep 7, 2026
Show 49 more resultsHide 49 results
- 21.63Sep 4, 2026
- 32.16Jun 18, 2026
- AA Intelligence13.00Jun 20, 2026Artificial Analysis Intelligence Index
- AA Intelligence13.31Oct 7, 2026aa_intelligence_index
- AA Intelligence24.74Oct 7, 2026aa_intelligence_index
- -34.45Oct 7, 2026
- 5.40Oct 7, 2026
- 94.17May 20, 2026
- 94.00Oct 6, 2026
- 94.00May 18, 2026
- 99.51Sep 7, 2026
- 72.83May 10, 2026
- 33.17May 10, 2026
- 5.83May 10, 2026
- 17.60Jul 29, 2026
- 1439Jul 23, 2026
- Artificial Analysis Coding Index49.39Sep 9, 2026aa_coding_index
- 27.30Jun 18, 2026
- 88.70Sep 7, 2026
- FrontierMath (overall)26.70Oct 6, 2026FrontierMath
- frontiermath_tier_4_v112.50Jun 20, 2026frontiermath_tier_4
- frontiermath_tier_4_v16.25May 20, 2026frontiermath_tier_4
- 95.00Jul 9, 2026
- 25.40Jul 9, 2026
- 50.90Jul 9, 2026
- 39.60Jul 9, 2026
- HLE (with tools)42.70May 18, 2026HLE (w/ Tools)
- 96.30May 18, 2026
- 91.67Jun 20, 2026
- Image input eval - extremism (not_unsafe)0.98Jul 9, 2026
- Image input eval - harms-erotic (not_unsafe)1.00Jul 9, 2026
- Image input eval - hate (not_unsafe)0.98Jul 9, 2026
- Image input eval - self-harm (not_unsafe)0.98Jul 9, 2026
- 76.88Sep 2, 2026
- 87.00May 18, 2026
- 91.00Jul 29, 2026
- 85.40Jul 29, 2026
- 48.33Sep 2, 2026
- 66.00Aug 29, 2026
- 67.00Oct 6, 2026
- 47.60Jul 29, 2026
- vectara_answer_rate100.00Aug 29, 2026Answer Rate
- vectara_avg_summary_length254.40Aug 29, 2026Average Summary Length (Words)
- vectara_factual_consistency87.90Aug 29, 2026Factual Consistency Rate
- 87.90Jun 20, 2026
- 82.70Jun 13, 2026
- τ²-Bench (Retail)77.90Oct 6, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)46.49Oct 7, 2026aa_tau2
- τ²-Bench Telecom (AA run)81.87Oct 7, 2026aa_tau2
GPT-5.1: common questions
Who makes GPT-5.1?
GPT-5.1 is made by OpenAI.
When was GPT-5.1 released?
GPT-5.1 was released on Nov 13, 2025, according to Artificial Analysis.
What is GPT-5.1 good at?
GPT-5.1 is capable in long context, instruction following, multimodal tasks, factuality, and coding; and behind the leaders in reasoning and agentic tasks. Too few results yet to rate safety, math, or multilingual tasks.
How much does GPT-5.1 cost?
GPT-5.1 costs $1.25 per million input tokens and $10.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 78% of the 329 priced models we track.
How many benchmarks has GPT-5.1 been tested on?
We track 105 results for GPT-5.1 on 66 benchmarks from 29 sources, 41 of them independently verified. The latest was recorded on Oct 7, 2026.
Which API features does GPT-5.1 support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.1.
About this record
Where GPT-5.1's numbers come from, and every name it appears under.
- Tracked since
- Jun 20, 2026
- Newest source mention
- Sep 7, 2026
Where the results come from
Verification: 105 scores · 41 independently verified · 34 aggregator-attributed · 15 vendor cross-reference · 15 vendor-reported. How these tiers are assigned
From 29 sources on 15 sites. Artificial Analysis supplies 36 of them; the 41 independently verified results come from 11 sites. Bars are coloured by trust tier.
- artificialanalysis.ai36
- huggingface.co9
- arcprize.org8
- deploymentsafety.openai.com8
- api.llm-stats.com7
- storage.googleapis.com6
- www-cdn.anthropic.com6
- datasets-server.huggingface.co5
- raw.githubusercontent.com5
- matharena.ai4
- epoch.ai3
- lmarena.ai3
- labs.scale.com2
- swebench.com2
- simple-bench.com1
Also known as
How our sources name GPT-5.1 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | — | gpt-5.1 (thinking, low) gpt-5.1-low-2025-11-13 |
| medium | GPT 5.1 (2025-11-13) (medium) | gpt-5.1 (2025-11-13) (medium reasoning) gpt-5.1 (medium) gpt-5.1 (thinking, medium) gpt-5.1-medium |
| high | — | gpt-5.1 (high) gpt-5.1 (thinking, high) gpt-5.1-high gpt-5.1-high-2025-11-13 |