GPT-5
GPT-5 is capable in long context, instruction following, and multimodal tasks; and behind the leaders in reasoning, factuality, coding, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Safety or Multilingual.
Price
$1.25input$10.00outputper million tokens
From Artificial Analysis · 2 providers tracked · All prices
Evidence
210results on111benchmarks
- 53 independently verified
- 65 aggregator
- 22 vendor-reported
- 70 cross-referenced
From 40 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
856 papers reference GPT-5GPT-5 benchmark results
210 results on 111 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
14.5% behind the leader1 of 3 ranked benchmarks measured
- 78.20Oct 8, 2026
18.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench73.06Oct 8, 2026aa_ifbench
- 69.60Oct 7, 2026
Show 6 more instruction following resultsHide 6 instruction following results
21.6% behind the leader3 of 6 ranked benchmarks measured
- 84.20Oct 7, 2026
- CharXiv (reasoning)81.10Oct 7, 2026CharXiv-R
- MMMU-Pro74.22Oct 8, 2026aa_mmmu_pro
Show 7 more multimodal resultsHide 7 multimodal results
27.3% behind the leader5 of 6 ranked benchmarks measured
- GPQA Diamond85.35Oct 8, 2026gpqa
- 56.70May 10, 2026
- Humanity's Last Exam28.50Oct 8, 2026aa_hle
- 5.71Oct 8, 2026
- 9.86May 10, 2026
Show 20 more reasoning resultsHide 20 reasoning results
- 0.00Sep 22, 2026
- 1.94May 10, 2026
- 7.49May 10, 2026
- 0.00Oct 8, 2026
- 1.14Oct 8, 2026
- GPQA Diamond67.27Oct 8, 2026gpqa
- GPQA Diamond80.81Oct 8, 2026gpqa
- GPQA Diamond84.18Oct 8, 2026gpqa
- GPQA Diamond87.30Oct 7, 2026GPQA
- GPQA Diamond85.70Oct 7, 2026GPQA
- 85.70Jul 13, 2026
- 85.00Jun 6, 2026
- Humanity's Last Exam6.02Oct 8, 2026aa_hle
- Humanity's Last Exam25.39Oct 8, 2026aa_hle
- Humanity's Last Exam19.56Oct 8, 2026aa_hle
- 24.80Oct 7, 2026
- Humanity's Last Exam24.80Jul 13, 2026HLE (no tools)
- Humanity's Last Exam26.00Jun 25, 2026HLE_text
- Humanity's Last Exam26.30Jun 12, 2026HLE (Text-only) no tools
- Humanity's Last Exam26.50Jun 6, 2026HLE (w/o tools)
29.5% behind the leader4 of 4 ranked benchmarks measured
- Vectara HHEM hallucination ratelower is better15.10Aug 10, 2026
- 50.10May 20, 2026
- AA-Omniscience · Accuracy40.32Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination17.82Oct 8, 2026omniscienceNonHallucination
Show 7 more factuality resultsHide 7 factuality results
- AA-Omniscience · Accuracy29.63Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy38.12Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy39.48Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination9.85Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination16.80Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination20.98Oct 8, 2026omniscienceNonHallucination
- Vectara HHEM hallucination ratelower is better14.70May 2, 2026
34.1% behind the leader8 of 10 ranked benchmarks measured
- 87.00May 18, 2026
- 74.90Oct 7, 2026
- SciCode42.94Sep 4, 2026aa_scicode
- 1417.82May 22, 2026
- SWE-bench Multilingual55.30Jun 12, 2026SWE-bench Multilingual w/ tools
- Terminal-Bench Hard32.58Oct 8, 2026aa_terminalbench_hard
- 41.78Oct 8, 2026
- Terminal-Bench 2.135.21Oct 8, 2026terminalbenchV21
Show 13 more coding resultsHide 13 coding results
- 84.50Jun 6, 2026
- SciCode38.77Sep 4, 2026aa_scicode
- SciCode39.12Sep 4, 2026aa_scicode
- SciCode41.09Sep 4, 2026aa_scicode
- 43.00Jun 6, 2026
- SciCode42.90May 15, 2026SciCode no tools
- 65.00Sep 25, 2026
- SWE-bench Verified74.90Jun 12, 2026SWE-bench Verified w/ tools
- Terminal-Bench Hard18.18Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard26.52Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench Hard37.88Oct 8, 2026aa_terminalbench_hard
- 31.00Jun 6, 2026
- 30.50May 18, 2026
35.5% behind the leader3 of 7 ranked benchmarks measured
- 54.90Oct 7, 2026
- τ-Bench V3 · Banking22.06Oct 8, 2026tauBanking
- 21.00Oct 8, 2026
Show 4 more agentic resultsHide 4 agentic results
- 54.90Jun 6, 2026
- 32.49Jun 15, 2026
- 25.07Jun 15, 2026
- 0.00Jun 15, 2026
41.2% behind the leader2 of 5 ranked benchmarks measured
- 55.44Sep 21, 2026
- 21.95Sep 21, 2026
Show 3 more math resultsHide 3 math results
- 18.25Sep 22, 2026
- 37.19Sep 21, 2026
- 76.00May 18, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 6.00Sep 22, 2026
- AA Intelligence23.00Sep 21, 2026Artificial Analysis Intelligence Index
- 81.30Sep 11, 2026
- 97.06Sep 7, 2026
- 99.75Sep 7, 2026
- 97.63Sep 7, 2026
Show 113 more resultsHide 113 results
- 26.53Sep 4, 2026
- 49.69Jun 18, 2026
- 45.83Jun 18, 2026
- 22.86Jun 18, 2026
- AA Intelligence35.00Jun 27, 2026Artificial Analysis Intelligence Index
- AA Intelligence11.45Oct 8, 2026aa_intelligence_index
- AA Intelligence22.98Oct 8, 2026aa_intelligence_index
- AA Intelligence20.79Oct 8, 2026aa_intelligence_index
- AA Intelligence22.87Oct 8, 2026aa_intelligence_index
- -33.80Oct 8, 2026
- -10.78Oct 8, 2026
- -10.87Oct 8, 2026
- -8.73Oct 8, 2026
- 86.70May 1, 2026
- 88.00Oct 7, 2026
- 95.00May 2, 2026
- 94.60Oct 7, 2026
- AIME 202594.60Jul 13, 2026AIME 2025 (no tools)
- AIME 202594.00Jun 6, 2026AIME25
- 100.00May 15, 2026
- 94.60Jun 12, 2026
- 99.60May 15, 2026
- 87.67Jun 20, 2026
- 99.10Sep 7, 2026
- 44.00May 10, 2026
- 56.17May 10, 2026
- 65.67May 10, 2026
- 92.20Jun 6, 2026
- 71.90Jun 6, 2026
- 73.00Jun 6, 2026
- Artificial Analysis Coding Index37.78Sep 9, 2026aa_coding_index
- 30.72Jun 18, 2026
- 38.95Jun 18, 2026
- 25.05Jun 18, 2026
- 96.80Sep 7, 2026
- 54.90Jun 12, 2026
- 72.89Jul 29, 2026
- 65.00Jun 6, 2026
- 63.00May 18, 2026
- 63.00Jun 12, 2026
- 65.70Oct 7, 2026
- 63.90Jun 6, 2026
- 48.50May 15, 2026
- 48.50Jun 12, 2026
- 86.00May 15, 2026
- 86.00Jun 12, 2026
- FrontierMath (overall)26.30Oct 7, 2026FrontierMath
- frontiermath_tier_4_v112.50Aug 29, 2026frontiermath_tier_4
- 76.40Jun 6, 2026
- 78.49May 22, 2026
- 64.78May 22, 2026
- 46.94May 22, 2026
- 66.11Jun 9, 2026
- GPQA (unspecified)85.70May 15, 2026GPQA no tools
- 67.20May 15, 2026
- 95.60Jul 9, 2026
- 34.70Jul 9, 2026
- 57.70Jul 9, 2026
- 67.20Jun 12, 2026
- 46.20Jul 9, 2026
- HLE (with tools)35.20Jun 6, 2026HLE (w/ tools)
- HLE (with tools)42.00May 15, 2026HLE (Text-only) heavy
- HLE (with tools)41.70May 15, 2026HLE (Text-only) w/ tools
- 93.30Oct 7, 2026
- HMMT 202593.30Jul 13, 2026HMMT 2025 (no tools)
- 93.20Jun 25, 2026
- 88.33May 11, 2026
- 88.30Jun 6, 2026
- 89.17May 10, 2026
- 89.20May 18, 2026
- 100.00May 15, 2026
- 93.30Jun 12, 2026
- 96.70May 15, 2026
- 93.40Oct 7, 2026
- 38.10Sep 2, 2026
- 20.83May 10, 2026
- 83.50Jun 25, 2026
- LiveCodeBench85.00Jun 6, 2026LiveCodeBench (LCB)
- 86.80Jul 13, 2026
- 87.00Jun 12, 2026
- Longform Writing eval (Kimi K2 Thinking system card)71.40Jun 12, 2026Longform Writing no tools
- 78.75Sep 2, 2026
- 68.44May 10, 2026
- 40.00May 10, 2026
- 100.00May 10, 2026
- 87.50Jun 6, 2026
- 87.00Jun 6, 2026
- MMLU-Pro87.10May 15, 2026MMLU-Pro no tools
- MMLU-Redux95.30May 15, 2026MMLU-Redux no tools
- Multi-SWE-Bench39.30Jun 12, 2026Multi-SWE-bench w/ tools
- 56.20May 15, 2026
- 56.20Jun 12, 2026
- 51.40May 15, 2026
- 51.40Jun 12, 2026
- 65.00Aug 10, 2026
- 62.60Oct 7, 2026
- 43.80Jun 6, 2026
- 35.20May 18, 2026
- 43.80Jun 12, 2026
- 99.90Aug 10, 2026
- 162.70Aug 10, 2026
- 109.70May 2, 2026
- 84.90Aug 10, 2026
- 85.30May 2, 2026
- 84.60Oct 7, 2026
- 77.80Jun 6, 2026
- 82.40Jun 13, 2026
- τ²-Bench85.00Jun 6, 2026τ²-Bench-Telecom
- τ²-Bench (Retail)81.10Oct 7, 2026Tau2 Retail
- τ²-Bench Telecom (AA run)66.96Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)84.80Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)84.21Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)86.55Oct 8, 2026aa_tau2
GPT-5: common questions
Who makes GPT-5?
GPT-5 is made by OpenAI.
When was GPT-5 released?
GPT-5 was released on Aug 7, 2025, according to Artificial Analysis.
What is GPT-5 good at?
GPT-5 is capable in long context, instruction following, and multimodal tasks; and behind the leaders in reasoning, factuality, coding, agentic tasks, and math. Too few results yet to rate safety or multilingual tasks.
How much does GPT-5 cost?
GPT-5 costs $1.25 per million input tokens and $10.00 per million output tokens, according to Artificial Analysis. We track its price at 2 providers. At a mix of three input tokens to one output token, it costs more than 78% of the 330 priced models we track.
How many benchmarks has GPT-5 been tested on?
We track 210 results for GPT-5 on 111 benchmarks from 40 sources, 53 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does GPT-5 support?
OpenRouter lists tool calling, structured outputs, and reasoning for GPT-5.
About this record
Where GPT-5's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Aug 24, 2026
Where the results come from
Verification: 210 scores · 53 independently verified · 65 aggregator-attributed · 70 vendor cross-reference · 22 vendor-reported. How these tiers are assigned
From 40 sources on 18 sites. Hugging Face supplies 68 of them; the 53 independently verified results come from 14 sites. Bars are coloured by trust tier.
- huggingface.co68
- artificialanalysis.ai67
- api.llm-stats.com18
- raw.githubusercontent.com10
- arcprize.org8
- epoch.ai6
- matharena.ai6
- storage.googleapis.com6
- x.ai5
- deploymentsafety.openai.com4
- aider.chat2
- datasets-server.huggingface.co2
- labs.scale.com2
- swebench.com2
- 99franklin.github.io1
- lmarena.ai1
- simple-bench.com1
- www-cdn.anthropic.com1
Also known as
How our sources name GPT-5 at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| minimal | — | gpt-5 (minimal) gpt-5-minimal gpt-5-minimal-2025-08-07 |
| low | — | gpt-5 (low) gpt-5-low |
| medium | GPT 5 (2025-08-07) (medium) | gpt-5 (2025-08-07) (medium reasoning) gpt-5 (medium) gpt-5-medium |
| high | gpt-5 high | gpt-5 (high) gpt-5 (high) (Non-Reasoning) gpt-5 (high) (reasoning) gpt-5-2025-08-07 (High) gpt-5-high-2025-08-07 |