NVIDIA’s Magpie TTS Multilingual is an open-weights, 364-million-parameter text-to-speech model that uses frame stacking — generating two audio frames per decoding step — plus local transformer layers to halve decoding iterations while preserving audio quality across twelve languages. Deployed via NVIDIA NIM containers on self-hosted infrastructure, it achieves 32-79ms time-to-first-audio latency depending on GPU hardware, enabling low-latency, privacy-preserving conversational voice agents with full data control.
