NVIDIA announced on Monday, August 24 that the Groq 3 LPX — its inference accelerator purpose-built for autonomous AI agent systems — has entered full production. Released as a companion to the Vera Rubin NVL72 platform, the chip marks the first time a dedicated ultra-fast token generation architecture is being made available for commercial use in large-scale data centers.
Unlike what is seen in conventional AI chips, which attempt to do everything with the same hardware, the Groq 3 LPX was conceived to solve a specific and chronic problem: decode latency. When an AI agent works with enormous context windows — 100,000 tokens or more — it must first process and understand that mass of information (the "prefill") and then generate its response token by token (the "decode"). Traditionally, GPUs have handled both stages. But while prefill can be massively parallelized, decode is inherently sequential. A single token must be generated before the next one can begin. This bottleneck causes agents to stall in visible pauses while processing complex chained tasks.
The Groq 3 LPX resolves this through a division of labor. Rubin GPUs process the context en masse; LPUs (language processing units, inherited from the December 2025 acquisition of Groq Inc.) take over token generation with deterministic latency. The result, according to NVIDIA, is a system where a complete 72-GPU rack works in tandem with up to 256 LP30 accelerators, interconnected through sixth-generation NVLink. This separation is not merely theoretical — the benchmark data is explicit. In Artificial Analysis tests, the Groq 3 LPX generated 3,400 tokens per second on the open-source Gemma 4 31B model with a 100,000-token context window. The company says this is four times more responsive than any competing platform, reducing tasks that previously took hours to minutes.
But the launch goes beyond benchmark numbers. NVIDIA already has customers. Nebius, an AI-focused cloud provider, was the first to adopt the chip and plans to integrate it into Nebius Token Factory, its production inference platform. According to Nebius CTO Danila Shtan, the goal is to ensure that "every step of an agent's loop feels instant." Groq — the very company whose chiplets gave rise to the LPU — will also be an adopter, as will SpaceX, which announced it will use the Vera CPU from the Vera Rubin platform to orchestrate agents across terrestrial data centers and orbital satellites.
What makes this architecture particularly significant is the shift in industry paradigm. For years, the dominant cost in AI was training — training a frontier model required clusters of thousands of GPUs running for months. Now, with increasingly capable models, the bottleneck has shifted to inference. A successful agent does not ask a single question; it executes hundreds or thousands of inference iterations per task — reading files, calling APIs, verifying results, rewriting code, and invoking external tools. Each iteration generates tokens. Each token costs infrastructure. If token generation is slow, the agent is not just slow — it is economically unviable.
This explains why NVIDIA is positioning Vera Rubin as a "token factory" — an architecture where compute, networking, and inference acceleration are designed as a unified system rather than stacked components. The communication fabric via Spectrum-X Multiplane, for example, can scale clusters of up to 512,000 GPUs in flat networks without the latency jitter that affects perceived performance. NVLink Fusion, meanwhile, allows custom CPUs and DPUs to connect directly to sixth-generation rack-scale systems.
The financial stakes are equally ambitious. The complete Vera Rubin ecosystem — including Rubin GPUs, Groq 3 LPUs, Vera CPUs, BlueField-4 DPUs, Spectrum-6 switches, and Scale-In software — is estimated to represent roughly $20 billion in data center infrastructure investment over the next three years. No other individual AI platform has scaled to that magnitude.
What is becoming clear is that the era of interactive inference has already begun. The Groq 3 LPX is not just another chip — it is confirmation that the AI market is migrating from training to inference as the primary engine of investment and innovation. As AI agents gain the ability to perform complex tasks in minutes rather than hours, a fundamental question emerges: how many enterprise processes will still be conducted by humans when machines can reason autonomously, verify results, and execute tools with near-zero perceived delay?
Sources: NVIDIA Newsroom, SiliconANGLE, NVIDIA Blog
✓ Independent sources cross-checked and verified before publishing