Nvidia has launched the Groq 3 LPX inference accelerator into full production, positioning it as a specialized chip designed to enhance AI agent responsiveness by dramatically increasing token generation speeds. The chip works alongside Nvidia’s Vera Rubin GPU platform to handle decode workloads, with Nvidia reporting benchmark speeds of 3,400 tokens per second. Nebius Group has become the first customer to commit to deploying the accelerators in its Token Factory production inference service. The technology was licensed from Groq Inc., which Nvidia acquired for $20 billion in December 2025.
