NVIDIA's Groq 3 LPX rack-scale inference system has entered full production, adding a dedicated decode tier to the Vera Rubin NVL72 platform for agent-heavy workloads.
NVIDIA says an Artificial Analysis benchmark running Gemma 4 31B, an open-source agentic model, logged 3,400 output tokens per second on 100,000-token contexts, a result NVIDIA puts at about four times the rate of the nearest alternative platform. The LPX was codesigned with Vera Rubin NVL72 to split the inference path: Rubin GPUs process large context windows while LPX accelerators absorb the token-by-token decode stage where agent latency compounds. Rack-scale deployments link LPX accelerators over direct chip-to-chip paths so fleets of LPUs act as one inference engine.
Reaching production matters because it moves the system from roadmap to orderable capacity: clouds can now commit real racks, which is what makes named adopters meaningful rather than symbolic. The timing also tracks a broader change in how AI infrastructure gets bought. Teams that spent recent cycles comparing training throughput now judge platforms on tokens per second and per-token economics, because that is where agent workloads spend their budgets.
Agent workloads make decode latency the binding constraint. A chatbot can hide a slow token in a longer pause, but an agent reasons, calls tools and hands results between systems, and each turn waits on the last token of the previous one. Long context makes it worse, because every step re-reads a large window, which is the regime where LPX is meant to earn its keep.
Named customers are already attached to the rollout. Nebius becomes the first AI cloud to adopt the LPX system, folding it into the Vera Rubin NVL72 racks of its Token Factory. CoreWeave has moved Spectrum-X Multiplane into production, joining Vera Rubin racks through parallel switch planes and skipping a third network tier. SpaceXAI plans to handle agent orchestration, tool use and code execution on Vera CPUs, from Earth-side data centers to orbital satellites.
The adopter list also previews NVIDIA's platform play: sell the AI factory as one codesigned system rather than separate chips, switches and cables. Spectrum-X Multiplane flattens the fabric between racks, and NVLink Fusion folds custom XPUs and CPUs into the same scale-up stack, letting hyperscalers blend NVIDIA silicon with their own designs.
The announcement carries no pricing for the LPX or per-token rates, so the cost advantage NVIDIA claims is still unverified, and the headline benchmark is NVIDIA's own. Independent measurement of long-context decode performance will decide whether the LPX pitch holds as clouds bring the racks online.













