Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

gpt-oss-120b

Score basis

gpt-oss-120b is behind the leaders in instruction following, long context, reasoning, and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.

Capability profile

Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.

Limited4 capabilities
  1. Instruction Following−33.0%2 of 3
  2. Long Context−36.0%1 of 3
  3. Reasoning−38.3%4 of 6
  4. Coding−43.9%4 of 10
Not rated6 capabilities

Too few results yet to rate Agentic, Safety, Math, Multimodal, Multilingual or Factuality.

Price

$0.15input$0.59outputper million tokens

From Artificial Analysis · 12 providers tracked · All prices

Cheaper than 77% of 330 priced models · 3:1 input-to-output blend, log scale

Evidence

62results on42benchmarks

  • 21 independently verified
  • 37 aggregator
  • 4 vendor-reported

From 19 sources · latest Oct 8, 2026 · How verification works

API features

Tool callingStructured outputsReasoning

As listed by OpenRouter

gpt-oss-120b benchmark results

62 results on 42 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.

Instruction Following

LimitedFull instruction following ranking

33.0% behind the leader2 of 3 ranked benchmarks measured

Show 1 more instruction following resultHide 1 instruction following result

Long Context

LimitedFull long context ranking

36.0% behind the leader1 of 3 ranked benchmarks measured

Show 1 more long context resultHide 1 long context result

Reasoning

LimitedFull reasoning ranking

38.3% behind the leader4 of 6 ranked benchmarks measured

Show 5 more reasoning resultsHide 5 reasoning results

Coding

LimitedFull coding ranking

43.9% behind the leader4 of 10 ranked benchmarks measured

Show 4 more coding resultsHide 4 coding results

Agentic

Not enough dataFull agentic ranking

0 of 7 ranked benchmarks measured

Show 4 more agentic resultsHide 4 agentic results

Factuality

Not enough dataFull factuality ranking

0 of 4 ranked benchmarks measured

Show 3 more factuality resultsHide 3 factuality results

More results

Benchmarks outside the capability baskets. They are not ranked against a leader.

Show 21 more resultsHide 21 results

gpt-oss-120b: common questions

Who makes gpt-oss-120b?

gpt-oss-120b is made by OpenAI.

When was gpt-oss-120b released?

gpt-oss-120b was released on Aug 5, 2025, according to Artificial Analysis.

What is gpt-oss-120b good at?

gpt-oss-120b is behind the leaders in instruction following, long context, reasoning, and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, multilingual tasks, or factuality.

How much does gpt-oss-120b cost?

gpt-oss-120b costs $0.15 per million input tokens and $0.59 per million output tokens, according to Artificial Analysis. We track its price at 12 providers. At a mix of three input tokens to one output token, it is cheaper than 77% of the 330 priced models we track.

How many benchmarks has gpt-oss-120b been tested on?

We track 62 results for gpt-oss-120b on 42 benchmarks from 19 sources, 21 of them independently verified. The latest was recorded on Oct 8, 2026.

Which API features does gpt-oss-120b support?

OpenRouter lists tool calling, structured outputs, and reasoning for gpt-oss-120b.

About this record

Where gpt-oss-120b's numbers come from, and every name it appears under.

Tracked since
May 1, 2026
Newest source mention
Oct 3, 2026

Where the results come from

Verification: 62 scores · 21 independently verified · 37 aggregator-attributed · 4 vendor-reported. How these tiers are assigned

From 19 sources on 10 sites. Artificial Analysis supplies 38 of them; the 21 independently verified results come from 9 sites. Bars are coloured by trust tier.

  • artificialanalysis.ai38
  • storage.googleapis.com6
  • api.llm-stats.com4
  • raw.githubusercontent.com4
  • matharena.ai3
  • labs.scale.com2
  • swebench.com2
  • aider.chat1
  • epoch.ai1
  • simple-bench.com1

Also known as

How our sources name gpt-oss-120b at each reasoning setting.

SettingShort formAPI id
low—gpt-oss-120B (low) gpt-oss-120b-low
highgpt oss 120b (high)—
Also listed asOpenAI GPT OSS 120Bgpt-oss:120bgpt-oss-120b:batchgpt-oss-120b-ultragpt-oss-120b-turboopenai/gpt-oss-120b