The SGLang team and NVIDIA achieved a 5x improvement in DeepSeek-V4 throughput on NVIDIA GB300 GPUs since the model’s Day-0 launch, reaching roughly 11,200 tokens per second per GPU at 50 tokens per second per user by June 2026, up from 2,200 tokens per second per GPU in April 2026. The gains came from kernel optimizations such as MHC fusion, KV Compression V2, and W4A4 MegaMoE, along with runtime improvements including better SWA budgeting and breakable CUDA graph support, plus stability fixes in SGLang and Dynamo. Similar throughput gains of 2.85x to 2.91x were observed on NVIDIA Blackwell Ultra hardware while maintaining production-relevant interactivity levels.
