Across 27 shared benchmarks, MiniMax 2.7 230B A10B scores higher on 4 and nvidia-nemotron-3-ultra-550b-a55b on 23. The widest gap is CritPt (no tools), where nvidia-nemotron-3-ultra-550b-a55b scores 3.1 against 0.6.
| Benchmark | MiniMax 2.7 230B A10B | nvidia-nemotron-3-ultra-550b-a55b |
|---|---|---|
| Airline | 75.3 | 81.5 |
| Apex-Shortlist (no tools) | 28.9 | 74.9 |
| Apex-Shortlist (with tools) | 51.9 | 84.8 |
| Banking | 14.6 | 22.6 |
| browsecomp | 54.1 | 44.4 |
| CritPt (no tools) | 0.6 | 3.1 |
| gdpval | 47.6 | 46.7 |
| GPQA Diamond | 86.6 | 87 |
| HLE | 23.1 | 26.7 |
| IFBench (prompt loose) | 74.6 | 81.7 |
| imo_answer_bench | 75.1 | 92.3 |
| LiveCodeBench v6 | 77.2 | 89 |
| MMLU-Pro | 81.9 | 86.8 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) | 78.4 | 83 |
| multichallenge | 42.5 | 63.8 |
| PinchBench | 77.6 | 90 |
| ProfBench (Search) | 52 | 56 |
| Retail | 84.9 | 86.4 |
| SciCode (subtask) | 38.3 | 44.6 |
| SWE-bench Multilingual | 71.8 | 67.7 |
| SWE-bench Verified | 75.3 | 70.7 |
| TauBench V3 - Average | 66.1 | 70.9 |
| Telecom | 89.6 | 92.9 |
| Terminal-Bench 2.1 | 55.5 | 56.4 |
| Vals.ai Financial Agent 1.1 - with web search | 50.5 | 53.7 |
| Vals.ai Financial Agent 1.1 - without web search | 51.3 | 60.1 |
| WMT24++ (en→xx) | 82.8 | 83.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.