gpt-oss-20b
gpt-oss-20b is behind the leaders in instruction following and reasoning. Too few results yet to rate coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or factuality.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Coding, Agentic, Safety, Long Context, Math, Multimodal, Multilingual or Factuality.
Price
$0.07input$0.18outputper million tokens
From Artificial Analysis · 7 providers tracked · All prices
Evidence
63results on39benchmarks
- 8 independently verified
- 34 aggregator
- 4 vendor-reported
- 17 cross-referenced
From 13 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Research
34 papers reference gpt-oss-20bgpt-oss-20b benchmark results
63 results on 39 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
31.7% behind the leader1 of 3 ranked benchmarks measured
- IFBench65.10Oct 8, 2026aa_ifbench
40.5% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond68.79Oct 8, 2026gpqa
- Humanity's Last Exam10.98Oct 8, 2026aa_hle
- 1.43Oct 8, 2026
Show 7 more reasoning resultsHide 7 reasoning results
- 0.00Oct 8, 2026
- GPQA Diamond61.11Oct 8, 2026gpqa
- GPQA Diamond71.50Oct 7, 2026GPQA
- GPQA Diamond71.46Aug 11, 2026GPQA Diamond (no tools)
- Humanity's Last Exam5.28Oct 8, 2026aa_hle
- 10.90Oct 7, 2026
- 10.90May 16, 2026
0 of 10 ranked benchmarks measured
- SciCode38.89Oct 8, 2026aa_scicode
- Terminal-Bench 2.113.86Oct 8, 2026terminalbenchV21
- Terminal-Bench Hard10.61Oct 8, 2026aa_terminalbench_hard
Show 8 more coding resultsHide 8 coding results
- LiveCodeBench v661.00Jun 13, 2026LCB v6
- SciCode34.03Sep 4, 2026aa_scicode
- 38.63Aug 11, 2026
- 41.93Aug 11, 2026
- 52.44Aug 11, 2026
- 34.00May 16, 2026
- 15.17Aug 11, 2026
- Terminal-Bench Hard4.55Oct 8, 2026aa_terminalbench_hard
0 of 7 ranked benchmarks measured
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
- τ-Bench V3 · Banking7.01Oct 8, 2026tauBanking
Show 3 more agentic resultsHide 3 agentic results
- 0.74Aug 12, 2026
- 28.30May 16, 2026
- 2.47Jun 15, 2026
0 of 3 ranked benchmarks measured
- 34.67Oct 8, 2026
- 31.00Oct 8, 2026
0 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy16.00Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Accuracy15.08Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination13.29Oct 8, 2026omniscienceNonHallucination
Show 1 more factuality resultHide 1 factuality result
- AA-Omniscience · Non-hallucination5.89Oct 8, 2026omniscienceNonHallucination
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- 100.00Sep 7, 2026
- 87.17Sep 7, 2026
- 99.65Sep 7, 2026
- 98.69Sep 7, 2026
- 96.70Sep 7, 2026
- 85.97May 19, 2026
Show 21 more resultsHide 21 results
- 1.40Sep 9, 2026
- 21.86Jun 18, 2026
- AA Intelligence8.97Oct 8, 2026aa_intelligence_index
- AA Intelligence9.95Oct 8, 2026aa_intelligence_index
- -63.05Oct 8, 2026
- -58.55Oct 8, 2026
- 89.17May 2, 2026
- AIME 202591.70May 16, 2026AIME 25
- Artificial Analysis Coding Index20.70Sep 9, 2026aa_coding_index
- 14.37Jun 18, 2026
- 71.50May 16, 2026
- 42.50Oct 7, 2026
- 13.76Aug 11, 2026
- 76.67May 11, 2026
- 86.73Aug 9, 2026
- 76.40Aug 11, 2026
- 57.20Aug 11, 2026
- TAU-bench (retail)54.80Oct 7, 2026TAU-bench Retail
- 47.70Jun 13, 2026
- τ²-Bench Telecom (AA run)60.23Oct 8, 2026aa_tau2
- τ²-Bench Telecom (AA run)50.29Oct 8, 2026aa_tau2
gpt-oss-20b: common questions
Who makes gpt-oss-20b?
gpt-oss-20b is made by OpenAI.
When was gpt-oss-20b released?
gpt-oss-20b was released on Aug 5, 2025, according to Artificial Analysis.
What is gpt-oss-20b good at?
gpt-oss-20b is behind the leaders in instruction following and reasoning. Too few results yet to rate coding, agentic tasks, safety, long context, math, multimodal tasks, multilingual tasks, or factuality.
How much does gpt-oss-20b cost?
gpt-oss-20b costs $0.07 per million input tokens and $0.18 per million output tokens, according to Artificial Analysis. We track its price at 7 providers. At a mix of three input tokens to one output token, it is cheaper than 91% of the 330 priced models we track.
How many benchmarks has gpt-oss-20b been tested on?
We track 63 results for gpt-oss-20b on 39 benchmarks from 13 sources, 8 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does gpt-oss-20b support?
OpenRouter lists tool calling, structured outputs, and reasoning for gpt-oss-20b.
About this record
Where gpt-oss-20b's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Aug 29, 2026
Where the results come from
Verification: 63 scores · 8 independently verified · 34 aggregator-attributed · 17 vendor cross-reference · 4 vendor-reported. How these tiers are assigned
From 13 sources on 5 sites. Artificial Analysis supplies 34 of them; the 8 independently verified results come from 2 sites. Bars are coloured by trust tier.
- artificialanalysis.ai34
- huggingface.co17
- storage.googleapis.com6
- api.llm-stats.com4
- matharena.ai2
Also known as
How our sources name gpt-oss-20b at each reasoning setting.
| Setting | Short form | API id |
|---|---|---|
| low | — | gpt-oss-20B (low) gpt-oss-20b-low |
| high | GPT OSS 20B (high) | — |