Nemotron 3 Super 120B A12B
Nemotron 3 Super 120B A12B is behind the leaders in instruction following, long context, reasoning, factuality, and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, or multilingual tasks.
Capability profile
Bars show the model's median result as a share of the leading model's, per capability. Select a row to see its results.
Too few results yet to rate Agentic, Safety, Math, Multimodal or Multilingual.
Price
$0.30input$0.90outputper million tokens
From Artificial Analysis · 3 providers tracked · All prices
Evidence
82results on68benchmarks
- 4 independently verified
- 20 aggregator
- 58 vendor-reported
From 7 sources · latest Oct 8, 2026 · How verification works
API features
Tool callingStructured outputsReasoning
As listed by OpenRouter
Nemotron 3 Super 120B A12B benchmark results
82 results on 68 benchmarks, grouped by capability. Each bar is the result as a share of the capability leader's; every result links to its source.
27.1% behind the leader2 of 3 ranked benchmarks measured
- IFBench71.50Oct 8, 2026aa_ifbench
- 55.23Oct 7, 2026
Show 1 more instruction following resultHide 1 instruction following result
- 72.56Oct 7, 2026
28.2% behind the leader1 of 3 ranked benchmarks measured
- 65.67Oct 8, 2026
Show 1 more long context resultHide 1 long context result
- 58.31Oct 7, 2026
31.8% behind the leader3 of 6 ranked benchmarks measured
- GPQA Diamond80.00Oct 8, 2026gpqa
- Humanity's Last Exam20.76Oct 8, 2026aa_hle
- 3.14Oct 8, 2026
Show 3 more reasoning resultsHide 3 reasoning results
- GPQA Diamond82.70Oct 7, 2026GPQA
- 22.82Oct 7, 2026
- Humanity's Last Exam18.50Jul 7, 2026HLE (no tools)
37.9% behind the leader2 of 4 ranked benchmarks measured
- AA-Omniscience · Accuracy24.32Oct 8, 2026omniscienceAccuracy
- AA-Omniscience · Non-hallucination13.04Oct 8, 2026omniscienceNonHallucination
41.6% behind the leader5 of 10 ranked benchmarks measured
- 53.73Oct 7, 2026
- SciCode36.23Oct 8, 2026aa_scicode
- 45.78Oct 7, 2026
- Terminal-Bench Hard28.79Oct 8, 2026aa_terminalbench_hard
- Terminal-Bench 2.138.58Oct 8, 2026terminalbenchV21
Show 1 more coding resultHide 1 coding result
- 42.05Oct 7, 2026
0 of 7 ranked benchmarks measured
- 1.84Oct 8, 2026
- 0.00Oct 8, 2026
- Terminal-Bench 4.00.00Oct 8, 2026
Show 3 more agentic resultsHide 3 agentic results
- 1.13Aug 12, 2026
- 31.28Oct 7, 2026
- τ-Bench V3 · Banking10.31Oct 8, 2026tauBanking
0 of 5 ranked benchmarks measured
- 84.85Sep 2, 2026
- 91.67Sep 2, 2026
- 85.19May 10, 2026
Show 1 more math resultHide 1 math result
- 91.67May 2, 2026
More results
Benchmarks outside the capability baskets. They are not ranked against a leader.
- τ²-Bench Telecom (AA run)67.84Oct 8, 2026aa_tau2
- -41.50Oct 8, 2026
- AA Intelligence12.83Oct 8, 2026aa_intelligence_index
- 4.13Sep 9, 2026
- Artificial Analysis Coding Index37.72Sep 9, 2026aa_coding_index
- 79.36Oct 7, 2026
Show 47 more resultsHide 47 results
- AGIEval-En77.92Aug 11, 2026AGIEval-EN (CoT)
- 90.21Oct 7, 2026
- AIME25 no tools92.20Jul 7, 2026AIME25 (no tools)
- 56.90Jul 7, 2026
- ARC-Challenge96.08Aug 11, 2026ARC-Challenge (25-shot)
- 73.88Oct 7, 2026
- 72.80Jul 7, 2026
- 85.72Aug 24, 2026
- GPQA Diamond (with tools)81.30Jul 7, 2026GPQA (with tools)
- GSM8K90.67Aug 11, 2026GSM8K (8-shot, CoT)
- 88.97Aug 11, 2026
- 94.73Oct 7, 2026
- 94.20Jul 7, 2026
- 95.50Jul 7, 2026
- 80.49Aug 11, 2026
- 73.40Jul 7, 2026
- 81.19Aug 23, 2026
- LiveCodeBench82.10Jul 7, 2026LiveCodeBench (v5 2024-07↔2024-12)
- MBPP81.71Aug 11, 2026MBPP (3-shot)
- Minerva Math84.84Aug 11, 2026Minerva Math (4-shot)
- 86.01Aug 11, 2026
- 83.73Oct 7, 2026
- MMLU-Pro74.43Aug 11, 2026MMLU-Pro (5-shot)
- 83.80Jul 7, 2026
- 79.50Jul 7, 2026
- 48.60Aug 11, 2026
- 85.47Aug 11, 2026
- PIQA (Acc.)83.90Aug 11, 2026PIQA (acc)
- 64.30Jul 7, 2026
- 91.75Oct 7, 2026
- 66.98Aug 11, 2026
- RULER 1M93.90Jul 7, 2026RULER @ 1M
- 83.03Aug 11, 2026
- RULER 256K96.70Jul 7, 2026RULER @ 256k
- RULER 512K95.70Jul 7, 2026RULER @ 512k
- 56.60Jul 7, 2026
- 42.30Jul 7, 2026
- 59.50Jul 7, 2026
- 56.25Oct 7, 2026
- 60.80Aug 24, 2026
- 61.10Jul 7, 2026
- 25.50Jul 7, 2026
- 25.78Sep 23, 2026
- 31.00Oct 7, 2026
- WinoGrande78.93Aug 11, 2026WinoGrande (5-shot)
- 86.80Jul 7, 2026
- τ²-Bench (Retail)62.83Oct 7, 2026Tau2 Retail
Nemotron 3 Super 120B A12B: common questions
Who makes Nemotron 3 Super 120B A12B?
Nemotron 3 Super 120B A12B is made by NVIDIA.
When was Nemotron 3 Super 120B A12B released?
Nemotron 3 Super 120B A12B was released on Mar 11, 2026, according to Artificial Analysis.
What is Nemotron 3 Super 120B A12B good at?
Nemotron 3 Super 120B A12B is behind the leaders in instruction following, long context, reasoning, factuality, and coding. Too few results yet to rate agentic tasks, safety, math, multimodal tasks, or multilingual tasks.
How much does Nemotron 3 Super 120B A12B cost?
Nemotron 3 Super 120B A12B costs $0.30 per million input tokens and $0.90 per million output tokens, according to Artificial Analysis. We track its price at 3 providers. At a mix of three input tokens to one output token, it is cheaper than 62% of the 330 priced models we track.
How many benchmarks has Nemotron 3 Super 120B A12B been tested on?
We track 82 results for Nemotron 3 Super 120B A12B on 68 benchmarks from 7 sources, 4 of them independently verified. The latest was recorded on Oct 8, 2026.
Which API features does Nemotron 3 Super 120B A12B support?
OpenRouter lists tool calling, structured outputs, and reasoning for Nemotron 3 Super 120B A12B.
About this record
Where Nemotron 3 Super 120B A12B's numbers come from, and every name it appears under.
- Tracked since
- Apr 27, 2026
- Newest source mention
- Sep 23, 2026
Where the results come from
Verification: 82 scores · 4 independently verified · 20 aggregator-attributed · 58 vendor-reported. How these tiers are assigned
From 7 sources on 4 sites. Hugging Face supplies 38 of them; the 4 independently verified results come from 1 site. Bars are coloured by trust tier.
- huggingface.co38
- api.llm-stats.com20
- artificialanalysis.ai20
- matharena.ai4