Across 29 shared benchmarks, GLM-5.1-744B-A40B scores higher on 15 and nvidia-nemotron-3-ultra-550b-a55b on 14. The widest gap is Banking, where nvidia-nemotron-3-ultra-550b-a55b scores 22.6 against 12.8.
| Benchmark | GLM-5.1-744B-A40B | nvidia-nemotron-3-ultra-550b-a55b |
|---|---|---|
| Airline | 85 | 81.5 |
| Apex-Shortlist (no tools) | 71.1 | 74.9 |
| Apex-Shortlist (with tools) | 79 | 84.8 |
| Banking | 12.8 | 22.6 |
| browsecomp | 59.4 | 44.4 |
| CritPt (no tools) | 3.7 | 3.1 |
| gdpval | 54.7 | 46.7 |
| GPQA Diamond | 86.1 | 87 |
| HLE | 27.2 | 26.7 |
| HLE (with tools) | 50.4 | 37.4 |
| IFBench (prompt loose) | 76.6 | 81.7 |
| imo_answer_bench | 91.1 | 92.3 |
| IOI 2025 | 456.5 | 570 |
| LiveCodeBench v6 | 85.7 | 89 |
| MMLU-Pro | 85.9 | 86.8 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) | 85.8 | 83 |
| multichallenge | 63 | 63.8 |
| PinchBench | 81.2 | 90 |
| ProfBench (Search) | 46 | 56 |
| Retail | 84.1 | 86.4 |
| SciCode (subtask) | 47.7 | 44.6 |
| SWE-bench Multilingual | 74.8 | 67.7 |
| SWE-bench Verified | 76.2 | 70.7 |
| TauBench V3 - Average | 69.7 | 70.9 |
| Telecom | 96.9 | 92.9 |
| Terminal-Bench 2.1 | 59.3 | 56.4 |
| Vals.ai Financial Agent 1.1 - with web search | 60.7 | 53.7 |
| Vals.ai Financial Agent 1.1 - without web search | 60.2 | 60.1 |
| WMT24++ (en→xx) | 84.4 | 83.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.