VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

NVIDIA Ships Nemotron 3.5 Lightning: 30B MoE Model Built for Agent Execution

NVIDIA releases Nemotron 3.5 Lightning, a 31.6B-parameter open MoE model with 3.6B active params, targeting agent execution with nearly 670 tokens/sec…

Vector Wire — AI-assisted editorial illustration

NVIDIA has released Nemotron 3.5 Lightning, an open-weights mixture-of-experts model with 31.6 billion total parameters and only 3.6 billion active at any given time, purpose-built for the high-volume execution layer of long-running AI agents1,2,3.

The model is the first in NVIDIA's new Nemotron 3.5 lineup and directly succeeds the Nemotron 3 Nano 30B A3B, retaining its hybrid Mamba-Transformer architecture. It is now available on Ollama and through serverless inference from DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe.

NVIDIA frames the model's target workload as tool calls, result validation, and subagent delegation — tasks where deploying a frontier reasoning model at every step adds unnecessary cost and latency. Nemotron 3.5 Lightning is designed to sit inside agent harnesses, handling execution while a larger orchestrator model handles planning.

On the Artificial Analysis Intelligence Index, the model scores 24, a nine-point jump from its predecessor's score of 15. That ties OpenAI's gpt-oss-120b, which also scores 24, and trails NVIDIA's own Nemotron 3 Super at 26 — a model roughly four times larger. Smaller models optimized for raw intelligence, such as Qwen3.6 35B A3B (32) and Meta's Muse Glimmer (35), hold a clear lead on that index.

The speed story is where NVIDIA is staking its claim. In pre-release tests using final NVFP4 weights, Nemotron 3.5 Lightning reached nearly 670 tokens per second, which the-decoder.com reports is almost twice the throughput of Google's Gemini 3.5 Flash-Lite at 386 tokens per second. On GDPval-AA v2, the model reaches an Elo rating of 824, a 334-point gain over Nemotron 3 Nano and ahead of both gpt-oss-120b at 800 and Nemotron 3 Super at 698. On Terminal-Bench v2.1, its score jumps from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent.

NVIDIA ships the model under the permissive OpenMDW-1.1 license with both BF16 and NVFP4 weight formats available. According to Artificial Analysis, the NVFP4 variant also scores 24 on the Intelligence Index with minimal quality loss compared to the higher-precision version. The model is text-only with a context window of one million tokens.

Artificial Analysis reports that NVIDIA worked with partners including CodeRabbit and Harvey on post-training to boost performance in specific domains.

ANALYSIS The positioning is deliberate: rather than competing on peak intelligence against larger frontier models, NVIDIA is optimizing for the throughput-per-watt economics of agent execution — the repetitive inner loop of tool use and validation that dominates runtime in production agent systems. Tying gpt-oss-120b on the Intelligence Index while nearly doubling its inference speed makes the cost-per-task argument concrete. The hybrid Mamba-Transformer architecture and aggressive parameter sparsity (3.6B active out of 31.6B total) are the technical levers enabling that tradeoff.

The breadth of serverless inference partners at launch — seven providers — signals that NVIDIA is treating this as an ecosystem play, not just a model drop. Coupling open weights under a permissive license with day-one cloud availability lowers the barrier for agent framework developers to adopt Lightning as a default execution-tier model.

CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →