Across 10 shared benchmarks, DeepSeek-V3.2-Exp-Base scores higher on 6 and NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 on 4. The widest gap is humaneval, where NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 scores 77.4 against 67.7.
| Benchmark | DeepSeek-V3.2-Exp-Base | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 |
|---|---|---|
| AGIEval-En | 70.1 | 70 |
| arc_challenge | 95.5 | 92.7 |
| GSM8K | 91.1 | 91.3 |
| hellaswag | 89.4 | 85.5 |
| humaneval | 67.7 | 77.4 |
| MBPP | 75.6 | 78.6 |
| mmlu | 87.8 | 78.6 |
| MMLU-Pro | 63.3 | 67.9 |
| OpenBookQA | 48.2 | 47.6 |
| winogrande | 85.6 | 80 |
Best tracked score per model per benchmark (default configuration; source-attributed). ↓ marks lower-is-better metrics. Open either model for its full surface, provenance and pricing.