MMA News

Nvidia Groq 3 LPX Racks to Launch This Year Following Major Acquisition

August 25, 2026Carlos Mendoza2 мин

Nvidia has announced that its Groq 3 LPX rack is now in full production. This marks the commercial release of technology stemming from the company's most significant acquisition to date. The Groq rack is slated for deployment alongside Vera central processors and Rubin graphics processors at Nebius, and is expected to be operational by the end of this year, according to Nvidia senior director Dion Harris.

The accelerated production of Groq's chip and its availability to customers underscore the increasing demand for low-latency inference. This is crucial for creating AI agents that respond quickly to user input, particularly in applications like coding, where prolonged delays are undesirable. Nvidia suggests that this capability allows cloud providers to offer higher-priced services.

"For those serving tokens, it opens up the possibility of offering premium service tiers to users and customers who require the most latency-sensitive service agreements," Harris stated during a call.

In December, Nvidia acquired assets from the chip startup Groq for approximately $20 billion, representing the company's largest purchase. The Groq architecture features 500 megabytes of high-speed SRAM integrated directly onto the chip to mitigate memory-related bottlenecks. Groq chips are manufactured by Samsung, while Nvidia's GPUs are produced by Taiwan Semiconductor Manufacturing.

Nvidia's Groq 3 LPX racks consist of 256 individual Groq 3 chips. Nvidia claims that these racks can process 3,400 tokens per second, based on benchmarks from Artificial Analysis.

This is a highly competitive market. Advanced Micro Devices (AMD), a smaller GPU manufacturer, recently announced its intention to integrate its rack-scale systems with chips from Cerebras, a company that recently went public and is also focused on low-latency inference. OpenAI's new "Ultrafast" mode promises 750 tokens per second and is reportedly "powered by Cerebras."

Low-latency chips are not intended to replace GPUs, which are the foundational components for AI tasks, capable of both training and inference. GPUs also offer the flexibility to adapt to new technologies and models. Low-latency chips like Groq primarily focus on the "decode" phase of model serving.

"This is not about replacing GPUs," Harris clarified. "It's about utilizing the right processor at the right price for the appropriate part of the workload."

Nvidia is currently increasing shipments of its Vera Rubin systems, which entered production earlier this year. At the unveiling of the Vera Rubin and Groq 3 LPX systems in March, Nvidia CEO Jensen Huang projected cumulative sales of $1 trillion between the current-generation Blackwell chips and the new Vera Rubin systems through 2027.

At that time, Huang indicated that a quarter of the data center space designated for coding applications would be allocated to Groq chips.

"The remainder of my data center is entirely Vera Rubin," Huang added.

Nvidia is scheduled to report its earnings on Wednesday.