Across 30 shared benchmarks, nvidia-nemotron-3-ultra-550b-a55b scores higher on 14 and Qwen 3.5 397B 17B on 16. The widest gap is Apex-Shortlist (with tools), where nvidia-nemotron-3-ultra-550b-a55b scores 84.8 against 60.4.
| Benchmark | nvidia-nemotron-3-ultra-550b-a55b | Qwen 3.5 397B 17B |
|---|---|---|
| Airline | 81.5 | 76.5 |
| Apex-Shortlist (no tools) | 74.9 | 61.4 |
| Apex-Shortlist (with tools) | 84.8 | 60.4 |
| Banking | 22.6 | 20.9 |
| browsecomp | 44.4 | 40.5 |
| CritPt (no tools) | 3.1 | 2.4 |
| gdpval | 46.7 | 34.6 |
| GPQA Diamond | 87 | 87.1 |
| HLE | 26.7 | 28.5 |
| HLE (with tools) | 37.4 | 48.3 |
| IFBench (prompt loose) | 81.7 | 78.2 |
| imo_answer_bench | 92.3 | 84.5 |
| IOI 2025 | 570 | 441.3 |
| LiveCodeBench v6 | 89 | 79.3 |
| longbench_v2 | 61.9 | 68.9 |
| MMLU-Pro | 86.8 | 88.3 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) | 83 | 86.4 |
| multichallenge | 63.8 | 63.9 |
| PinchBench | 90 | 86.6 |
| ProfBench (Search) | 56 | 53 |
| Retail | 86.4 | 88.5 |
| SciCode (subtask) | 44.6 | 48 |
| SWE-bench Multilingual | 67.7 | 70.9 |
| SWE-bench Verified | 70.7 | 73.6 |
| TauBench V3 - Average | 70.9 | 71 |
| Telecom | 92.9 | 98 |
| Terminal-Bench 2.1 | 56.4 | 49.9 |
| Vals.ai Financial Agent 1.1 - with web search | 53.7 | 59 |
| Vals.ai Financial Agent 1.1 - without web search | 60.1 | 61.3 |
| WMT24++ (en→xx) | 83.7 | 86.8 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.