NVIDIA announced that Groq 3 LPX, its interactive AI inference accelerator, is now in full production, delivering world-class token generation speeds for agentic AI workloads at the Hot Chips conference on Sunday.
The accelerator, built for the Vera Rubin platform, achieved 3,431 output tokens per second on Gemma 4 31B with 100,000 tokens of input context, according to benchmarking by Artificial Analysis. The device is designed to accelerate the decode phase of inference, the part users experience as latency while an AI agent thinks, calls tools, and streams responses.
Nebius Leads Deployment
Nebius will be the first AI cloud to bring Groq 3 LPX to production via its Token Factory platform. CTO Danila Shtan said the goal is to make every step of an agents loop feel instant through the same API developers already use, with no migration to a new stack. Following Nebius, the purpose-built inference cloud Groq plans to be among the earliest adopters.
The LPX extends the Vera Rubin NVL72 platform by adding LPU-based decode acceleration alongside GPUs for prefill, creating a heterogeneous architecture. NVIDIA says the combination of Rubin GPUs for prefill and Groq LPU for fast decode addresses the reality that one chip topology does not win both batch economics and interactive decode.
Implications for Agentic AI
The launch targets a specific pain point: interactive AI applications where decode latency directly determines user experience. Coding agents, live collaboration tools, and long-context applications that need real-time streaming benefit most from the accelerated token generation.
NVIDIA has not published per-token pricing for LPX. Access is expected through Nebius Token Factory tier selection, with premium pricing versus standard GPU tiers. Analysts note that the comparison should focus on cost per successful agent task rather than cost per million tokens, since faster decode may reduce user abandonment and improve task completion rates.
The launch sits in the broader context of NVIDIA expanding beyond hardware into the AI access layer, following its $20 billion acquisition of Groqs assets and its recent $6 billion commitment to Poolside for open-weight model development.
discussion