Nvidia has moved Groq 3 LPX into full production, positioning the inference system as a fast token generator alongside its Vera Rubin rack-scale platform. The combination targets AI agents that process long contexts, call tools and keep generating output over many steps.
In a benchmark cited by Nvidia, Groq 3 LPX produced 3,400 output tokens per second on the Gemma 4 31B model with a 100,000-token context. The company says that was four times faster than the nearest alternative platform in the test. Nebius is the first AI cloud provider announced for the product, while CoreWeave has deployed Spectrum-X Multiplane networking for Vera Rubin racks. SpaceXAI also plans to use Vera CPUs.
Those announcements show where Nvidia expects inference demand to grow: from short chatbot replies toward persistent agent workflows. They do not establish the same advantage across all models or serving conditions. Developers will need cloud availability, pricing and independent workload tests before deciding whether the specialized generation layer improves their own applications.