Across 29 shared benchmarks, Kimi K2.6 1T A32B scores higher on 23 and nvidia-nemotron-3-ultra-550b-a55b on 5, with 1 level. The widest gap is CritPt (no tools), where Kimi K2.6 1T A32B scores 9.1 against 3.1.
| Benchmark | Kimi K2.6 1T A32B | nvidia-nemotron-3-ultra-550b-a55b |
|---|---|---|
| Airline | 85.8 | 81.5 |
| Apex-Shortlist (no tools) | 77.4 | 74.9 |
| Apex-Shortlist (with tools) | 73.2 | 84.8 |
| Banking | 23.1 | 22.6 |
| browsecomp | 61.3 | 44.4 |
| CritPt (no tools) | 9.1 | 3.1 |
| gdpval | 50.4 | 46.7 |
| GPQA Diamond | 91 | 87 |
| HLE | 34.8 | 26.7 |
| HLE (with tools) | 54 | 37.4 |
| IFBench (prompt loose) | 73.7 | 81.7 |
| imo_answer_bench | 93.7 | 92.3 |
| IOI 2025 | 585 | 570 |
| LiveCodeBench v6 | 90.2 | 89 |
| MMLU-Pro | 88.1 | 86.8 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) | 85 | 83 |
| multichallenge | 63.1 | 63.8 |
| PinchBench | 90.2 | 90 |
| ProfBench (Search) | 56 | 56 |
| Retail | 82.9 | 86.4 |
| SciCode (subtask) | 52 | 44.6 |
| SWE-bench Multilingual | 77.1 | 67.7 |
| SWE-bench Verified | 75.7 | 70.7 |
| TauBench V3 - Average | 72.4 | 70.9 |
| Telecom | 97.8 | 92.9 |
| Terminal-Bench 2.1 | 67.2 | 56.4 |
| Vals.ai Financial Agent 1.1 - with web search | 58.8 | 53.7 |
| Vals.ai Financial Agent 1.1 - without web search | 54 | 60.1 |
| WMT24++ (en→xx) | 84.5 | 83.7 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.