Across 27 shared benchmarks, Nemotron 3 Ultra 550B A55B scores higher on 1 and nvidia-nemotron-3-ultra-550b-a55b on 0, with 26 level. The widest gap is HLE, where Nemotron 3 Ultra 550B A55B scores 28.4 against 26.7. Both cost $0.60 per million input tokens on tracked API pricing; nvidia-nemotron-3-ultra-550b-a55b is cheaper on output ($2.75 against $3.60).
| Benchmark | Nemotron 3 Ultra 550B A55B | nvidia-nemotron-3-ultra-550b-a55b |
|---|---|---|
| Apex-Shortlist (no tools) | 74.9 | 74.9 |
| Apex-Shortlist (with tools) | 84.8 | 84.8 |
| browsecomp | 44.4 | 44.4 |
| CritPt (no tools) | 3.1 | 3.1 |
| gdpval | 46.7 | 46.7 |
| GPQA Diamond | 87 | 87 |
| HLE | 28.4 | 26.7 |
| HLE (with tools) | 37.4 | 37.4 |
| IFBench (prompt loose) | 81.7 | 81.7 |
| imo_answer_bench | 92.3 | 92.3 |
| IOI 2025 | 570 | 570 |
| LiveCodeBench v6 | 89 | 89 |
| longbench_v2 | 61.9 | 61.9 |
| MMLU-Pro | 86.8 | 86.8 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) | 83 | 83 |
| multichallenge | 63.8 | 63.8 |
| PinchBench | 90 | 90 |
| ProfBench (Search) | 56 | 56 |
| ruler | 94.7 | 94.7 |
| SciCode (subtask) | 44.6 | 44.6 |
| SWE-bench Multilingual | 67.7 | 67.7 |
| SWE-bench Verified | 70.7 | 70.7 |
| TauBench V3 - Average | 70.9 | 70.9 |
| Terminal-Bench 2.1 | 56.4 | 56.4 |
| Vals.ai Financial Agent 1.1 - with web search | 53.7 | 53.7 |
| Vals.ai Financial Agent 1.1 - without web search | 60.1 | 60.1 |
| WMT24++ (en→xx) | 83.7 | 83.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing. Quoted rates are the price-setter row we currently track for each model — its direct or vendor-official listing where one exists (nvidia, direct), otherwise the lowest tracked offer.