Full Breakdown
Nvidia’s Groq 3 LPX Accelerator Enters Full Production, Targeting Agentic AI Workloads
8/25/2026, 7:55:57 AM
Core Event
On August 24 2026, Nvidia announced that its Groq 3 LPX inference accelerator has entered full-volume production. The rack-scale system, which integrates 256 low-latency processing units (LPUs) per rack, is slated for deployment later in 2026. Nebius, an AI cloud provider, will be the first to field the hardware through its “Nebius Token Factory” platform.
Background & Context
Nvidia’s move into low-latency inference follows its December 2025 $20 billion licensing deal with Groq Inc., bringing Groq’s founder Jonathan Ross, president Sunny Madra, and much of the engineering team into Nvidia’s ecosystem while keeping Groq independent. The Groq 3 LPX pairs this technology with Nvidia’s Vera Rubin NVL72 platform to serve “agentic” AI systems that perform multi-step tasks.
Key Figures & Groups
- Jensen Huang – Founder and CEO of Nvidia.
- Dion Harris – Senior Director, Nvidia.
- Danila Shtan – CTO of Nebius.
- Jonathan Ross – Founder and CEO of Groq.
Data & Statistics
- Each rack houses 256 LPUs, delivering 3,400 output tokens per second on the Gemma 4 31B model with a 100,000-token context window.
- Nvidia claims the system is four times faster than the nearest competing platform for latency-sensitive workloads.
- Combined with Vera Rubin GPUs, the architecture can achieve up to 30× higher throughput per megawatt and 35× lower token cost on agentic workloads, according to internal tests.
- LPUs feature 500 MB of on-die SRAM and roughly 150 TB/s memory bandwidth per chip; each rack supplies 128 GB of high-bandwidth SRAM overall.
- The accelerator is manufactured by Samsung; Nvidia’s GPUs continue to be sourced from TSMC.
Official Statements & Responses
- Dion Harris told reporters the Groq 3 LPX will be deployed alongside Vera Rubin and Vera CPU racks in Nebius data centers, emphasizing complementarity with GPUs.
Why It Matters / Impact
The launch signals a shift toward inference efficiency for agentic AI that must generate thousands of tokens across many reasoning steps. Offloading the latency-critical decode phase to specialized LPUs lowers power consumption and improves user-perceived responsiveness, addressing data-center power constraints that limit scaling of AI workloads.
Conflicting Reports & Gaps
Nvidia has not disclosed shipment volumes, pricing, or the exact cloud-deployment model. Benchmark details beyond the 3,400 tokens/second figure—such as performance on larger mixture-of-expert models—remain unverified, and the company has not clarified handling of models larger than 31 billion parameters.
Verbatim Quotes
- “Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” — Jensen Huang
- “Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” — Danila Shtan
What’s Next
Nvidia indicated earnings will be reported later this week, with expected guidance on revenue contributions from the Groq 3 LPX line. Additional cloud partners, including Groq’s own inference service, are slated to adopt the accelerator in the coming months.
