- Sources: primary, discussion
- Summary: NVIDIA released Nemotron 3.5 Lightning, which the model card summarizes as a Mamba-2, mixture of experts, and attention hybrid, at 30B parameters with 3B active. The model card gives a release date of 2026-08-11 and states the NVFP4 build deploys on a single H100 or a single DGX Spark GB10 at up to 1M tokens of context, under the OpenMDW 1.1 license. The benchmark and speculative-decoding claims on the card are vendor-run and not independently reproduced.
- Why it matters: A 30B-class agent model with a long context fits on one desk-side machine.
- Follow-up: Whether an independent party reproduces the throughput and benchmark numbers.
send feedback on this story