Menu
Close
Achmadnurhidayat.id

News You Trust

Nvidia Begins Full Production of Groq 3 LPX Chip

Smallest Font
Largest Font

Nvidia Corp. announced on August 24, 2026, at the Hot Chips conference that its Groq 3 LPX artificial intelligence inference accelerator has entered full production as a purpose-built extension to the flagship Vera Rubin data center platform. The accelerator targets low-latency decode workloads and ultrafast token generation for agentic artificial intelligence systems, which require real-time continuous processing loops for coding, reasoning, and tool execution. The commercial rollout follows Nvidia's $20 billion asset acquisition of chip startup Groq in December, marking the largest acquisition in the semiconductor giant's history.

In benchmark testing conducted by Artificial Analysis, the Groq 3 LPX system generated 3,400 output tokens per second on Google's open-source Gemma 4 31B model using a 100,000-token input sequence, delivering four times the responsiveness of competing platforms.

Neocloud provider Nebius Group N.V. became the first cloud customer to commit to the hardware, planning to deploy it within its Nebius Token Factory production inference platform later this year. Nvidia executive Jensen Huang outlined the strategic vision for the Vera Rubin platform during its hardware expansion announcement.

"Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency," said Jensen Huang, founder and CEO of NVIDIA.

The chief executive emphasized that the architecture is optimized specifically to handle the demanding processing requirements of next-generation autonomous systems.

"Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide," said Jensen Huang, founder and CEO of NVIDIA.

Addressing the hardware setup, senior director Dion Harris clarified how the specialized accelerators fit alongside general-purpose graphics processing units in high-density data centers.

"For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive," said Dion Harris, senior director at Nvidia.

Harris explained that the LPX system is designed to offload memory-intensive decode phases without replacing primary GPU compute power.

"This isn't about replacing GPUs," said Dion Harris, senior director at Nvidia.

The executive highlighted that the complementary architecture balances cost and hardware capabilities for data center operators.

"It's about using the right price, right processor for the right part of the workload," said Dion Harris, senior director at Nvidia.

From the cloud provider side, Nebius executive Danila Shtan detailed how the hardware integration will affect end-user applications on their platform.

"Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate," said Danila Shtan, chief technology officer of Nebius.

Shtan noted that the deployment allows developers to leverage high-speed processing without revising their existing software integration layers.

"As the first AI cloud bringing it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant — through the same API developers are already using, with no migration to a new stack," said Danila Shtan, chief technology officer of Nebius.

Each LPX rack packages 256 individual Groq 3 chips manufactured by Samsung, featuring 500 megabytes of high-speed on-die SRAM per chip to bypass traditional memory bandwidth bottlenecks.

Nvidia also revealed that SpaceX has signed on to deploy Vera Rubin architecture and Vera CPUs across its ground data centers and satellite network.

Follow achmadnurhidayat.id Add to preferred sources on Google
Editors Team
Daisy Floren

What's Your Reaction?

  • Like
    0
    Like
  • Dislike
    0
    Dislike
  • Funny
    0
    Funny
  • Angry
    0
    Angry
  • Sad
    0
    Sad
  • Wow
    0
    Wow