gpt-oss-120b
gpt-oss-120b is behind the leaders in instruction following, long context, reasoning, and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Math, Multimodal, Multilingual or Factuality.
Price
$0.15input$0.59outputper million tokens
From Artificial Analysis · 12 providers tracked · All prices
Evidence
62results on42benchmarks
- 21 independently verified
- 37 aggregator
- 4 vendor-reported
From 19 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
62 papers reference gpt-oss-120bgpt-oss-120b benchmark results
62 results on 42 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
33.0% behind the leader2 of 3 ranked benchmarks measured
- IFBench68.98Oct 8, 2026aa_ifbench
- 45.34Oct 8, 2026
Show 1 more instruction following resultHide 1 instruction following result
- IFBench58.30Oct 8, 2026aa_ifbench
36.0% behind the leader1 of 3 ranked benchmarks measured
- 52.00Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 46.00Oct 8, 2026
38.3% behind the leader4 of 6 ranked benchmarks measured
- GPQA Diamond78.18Oct 8, 2026gpqa
- Humanity's Last Exam19.60Oct 8, 2026aa_hle
- 22.10May 10, 2026
- 1.14Oct 8, 2026
Show 5 more reasoning resultsHide 5 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond67.17Oct 8, 2026gpqa
- GPQA Diamond80.10Oct 7, 2026GPQA
- Humanity's Last Exam5.89Oct 8, 2026aa_hle
- 14.90Oct 7, 2026
43.9% behind the leader4 of 10 ranked benchmarks measured
- SciCode34.03Oct 8, 2026aa_scicode
- Terminal-Bench Hard23.48Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.126.22Oct 8, 2026terminalbenchV21
- 16.20Oct 8, 2026
Show 4 more coding resultsHide 4 coding results
- SciCode36.00Sep 4, 2026aa_scicode
- 26.00Jul 29, 2026
- Terminal-Bench 2.113.86Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard5.30Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 3.10Oct 8, 2026
- 5.65Oct 8, 2026
- 5.57Oct 8, 2026
Show 4 more agentic resultsHide 4 agentic results
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking12.78Oct 8, 2026tauBanking
- τ-Bench V3 · Banking2.89Oct 8, 2026tauBanking
0 of 4 ranked benchmarks measured
- 13.90May 20, 2026
- 14.20May 2, 2026
- AA-Omniscience · Accuracy21.78Oct 8, 2026omniscienceAccuracy
Show 3 more factuality resultsHide 3 factuality results
- AA-Omniscience · Accuracy19.80Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination9.18Oct 8, 2026omniscienceNonHallucination
- AA-Omniscience · Non-hallucination8.60Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 41.80Aug 29, 2026
- AA Intelligence24.00Jun 24, 2026Artificial Analysis Intelligence Index
- 96.67Jun 23, 2026
- 88.00May 19, 2026
- 92.67May 10, 2026
- 100.00May 10, 2026
Show 21 more resultsHide 21 results
- 6.21Sep 9, 2026
- 0.96Aug 18, 2026
- AA Intelligence11.60Oct 8, 2026aa_intelligence_index
- AA Intelligence10.21Oct 8, 2026aa_intelligence_index
- -53.50Oct 8, 2026
- -49.25Oct 8, 2026
- 90.00May 2, 2026
- 99.49May 10, 2026
- Artificial Analysis Coding Index30.44Sep 9, 2026aa_coding_index
- Artificial Analysis Coding Index21.24Sep 9, 2026aa_coding_index
- 98.50May 10, 2026
- 100.00May 10, 2026
- 57.60Oct 7, 2026
- 90.00May 10, 2026
- 26.00May 1, 2026
- TAU-bench (retail)67.80Oct 7, 2026TAU-bench Retail
- vectara_answer_rate99.90May 2, 2026Answer Rate
- vectara_avg_summary_length135.20May 2, 2026Average Summary Length (Words)
- vectara_factual_consistency85.80May 2, 2026Factual Consistency Rate
- τ²-Bench Telecom (AA run)65.79Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)45.03Oct 8, 2026aa_tau2
gpt-oss-120b: common questions
Who makes gpt-oss-120b?
gpt-oss-120b is made by OpenAI.
When was gpt-oss-120b released?
gpt-oss-120b was released on Aug 5, 2025, according to Artificial Analysis.
What is gpt-oss-120b good at?
gpt-oss-120b is behind the leaders in instruction following, long context, reasoning, and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.
How much does gpt-oss-120b cost?
gpt-oss-120b costs $0.15 per million input tokens and $0.59 per million output tokens, according to Artificial Analysis. We track its price at 12 providers. At a mix of three input tokens to one output token, it is cheaper than 77% of the 330 priced models we track.
How many benchmarks has gpt-oss-120b been tested on?
We track 62 results for gpt-oss-120b on 42 benchmarks from 19 sources, 21 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does gpt-oss-120b support?
OpenRouter lists tool calling, structured outputs, and reasoning for gpt-oss-120b.
About this record
Where gpt-oss-120b's numbers come from, and every name it appears under.
- Tracked since
- May 1, 2026
- Newest source mention
- Oct 3, 2026
Where the results come from
Verification: 62 scores · 21 independently verified · 37 aggregator-attributed · 4 vendor-reported. How these tiers are assigned
From 19 sources on 10 sites. Artificial Analysis supplies 38 of them; the 21 independently verified results come from 9 sites. Bars are coloured by trust tier.
- artificialanalysis.ai38
- storage.googleapis.com6
- api.llm-stats.com4
- raw.githubusercontent.com4
- matharena.ai3
- labs.scale.com2
- swebench.com2
- aider.chat1
- epoch.ai1
- simple-bench.com1
Also known as
How our sources name gpt-oss-120b at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | — | gpt-oss-120B (low) gpt-oss-120b-low |
| high | gpt oss 120b (high) | — |