Nvidia Begins Full Production of Groq 3 LPX Inference Rack
Nvidia announced on August 24, 2026, that its Groq 3 LPX interactive AI inference accelerator rack has entered full production, expanding the company's Vera Rubin computing platform for latency-sensitive agentic AI workloads.
The deployment marks the commercialization of technology acquired through Nvidia's $20 billion purchase of Groq assets in December. Cloud provider Nebius will be the first to adopt the new hardware later this year alongside Vera central processors and Rubin graphics processors.
The Groq 3 LPX rack packages 256 individual Groq 3 chips, featuring 500 megabytes of on-die SRAM to eliminate memory bottlenecks. Manufactured by Samsung, the system reached a benchmarked rate of 3,400 output tokens per second on the Gemma 4 31B model with a 100,000-token context.
"Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency," said Jensen Huang, founder and CEO of Nvidia.
Huang stated that the new architecture addresses the processing requirements of agentic systems operating with high context demands.
"Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide," Huang said.
Low-latency inference technology targets the decode phase of model execution rather than replacing general GPUs, allowing operators to charge higher rates for latency-sensitive enterprise service agreements.
"For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive," said Dion Harris, senior director at Nvidia.
Harris clarified the hardware integration strategy within data center environments.
"This isn't about replacing GPUs. It's about using the right price, right processor for the right part of the workload," Harris said.
Concurrently, SpaceXAI announced plans to integrate Nvidia Vera CPUs into its terrestrial data centers for Grok and deploy an optimized Vera Rubin NVL72 system aboard its first-generation Starmind AI satellite in orbit.
"Agentic AI requires a new kind of computing system — one built not only to generate answers, but to take action," said Ian Buck, vice president of hyperscale and high-performance computing at Nvidia.
Buck highlighted the shift from ground-based data centers toward space-based computing infrastructure.
"Vera gives AI agents the CPU performance to act in real time — executing code, processing data and coordinating complex tasks. SpaceXAI is taking this architecture from massive AI factories to the next frontier of computing in orbit," Buck said.
The Vera CPU features 88 Olympus cores and high-bandwidth LPDDR5X memory, delivering up to 1.2TB/s of bandwidth to manage orchestration tasks.
"Vera gives us the CPU performance and memory bandwidth to run enormous amounts of orchestration, code and data processing while keeping GPUs doing what they do best," said Mike Nicolls, president of SpaceXAI.
Nicolls noted that the setup increases performance efficiency per watt of compute.
"That means higher-performance AI agents and more useful work from every watt of compute," Nicolls said.
The announcements coincide with the release of AgentX 1.0, an open-source multi-turn agentic benchmark running on 2 megawatts of compute to establish standardized industry metrics across hardware platforms.
What's Your Reaction?
-
0
Like -
0
Dislike -
0
Funny -
0
Angry -
0
Sad -
0
Wow