VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Questions this page answers

Which is better, GLM-5.1-744B-A40B or nvidia-nemotron-3-ultra-550b-a55b?
Across 29 shared benchmarks, GLM-5.1-744B-A40B scores higher on 15 and nvidia-nemotron-3-ultra-550b-a55b on 14. The widest gap is Banking, where nvidia-nemotron-3-ultra-550b-a55b scores 22.6 against 12.8.

GLM-5.1-744B-A40B vs nvidia-nemotron-3-ultra-550b-a55b

Across 29 shared benchmarks, GLM-5.1-744B-A40B scores higher on 15 and nvidia-nemotron-3-ultra-550b-a55b on 14. The widest gap is Banking, where nvidia-nemotron-3-ultra-550b-a55b scores 22.6 against 12.8.

Z.aivsNVIDIA29 shared benchmarks1514 head-to-head
BenchmarkGLM-5.1-744B-A40Bnvidia-nemotron-3-ultra-550b-a55b
Airline8581.5
Apex-Shortlist (no tools)71.174.9
Apex-Shortlist (with tools)7984.8
Banking12.822.6
browsecomp59.444.4
CritPt (no tools)3.73.1
gdpval54.746.7
GPQA Diamond86.187
HLE27.226.7
HLE (with tools)50.437.4
IFBench (prompt loose)76.681.7
imo_answer_bench91.192.3
IOI 2025456.5570
LiveCodeBench v685.789
MMLU-Pro85.986.8
MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko)85.883
multichallenge6363.8
PinchBench81.290
ProfBench (Search)4656
Retail84.186.4
SciCode (subtask)47.744.6
SWE-bench Multilingual74.867.7
SWE-bench Verified76.270.7
TauBench V3 - Average69.770.9
Telecom96.992.9
Terminal-Bench 2.159.356.4
Vals.ai Financial Agent 1.1 - with web search60.753.7
Vals.ai Financial Agent 1.1 - without web search60.260.1
WMT24++ (en→xx)84.483.7

Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.