vLLM now provides Day-0 support for NVIDIA’s Nemotron 3.5 Lightning, a customizable 30-billion-parameter hybrid mixture-of-experts model with only 3 billion active parameters, designed for agent tasks. The model offers up to 4x higher throughput than similarly sized open models through its hybrid MoE design and multi-token prediction. vLLM enables deployment of Nemotron 3.5 Lightning across NVIDIA platforms from edge devices to data centers via an OpenAI-compatible API, with support for speculative decoding techniques including MTP, DFlash, and DSpark.
