Mastodon Skip to content Breaking US C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboard
pulseofnations. Real News. Global Impact.
Subscribe
live markets
S&P 5007,677.28▲ 3.58%NASDAQ26,151.30▲ 4.71%DOW53,577.40▲ 3.14%GOLD4,716.00▲ 15.94%WTI80.40▼ 9.98%BRENT85.31▼ 11.85%EUR/USD1.1675▲ 2.62%USD/JPY158.95▼ 2.98%DXY98.91▼ 2.53%BTC$78,921▼ 2.45%ETH$2,462▼ 2.22%SOL$97.16▼ 4.70%TOTAL CRYPTO$2.67T▼ 3.80%

NVIDIA Groq 3 LPX Enters Production at Record Speed

NVIDIA launches Groq 3 LPX inference accelerator in full production, hitting 3,400 tokens per second on 100K context, with Nebius as first cloud adopter.

Partner Surfshark VPN

NVIDIA announced that Groq 3 LPX, its interactive AI inference accelerator, is now in full production, delivering world-class token generation speeds for agentic AI workloads at the Hot Chips conference on Sunday.

The accelerator, built for the Vera Rubin platform, achieved 3,431 output tokens per second on Gemma 4 31B with 100,000 tokens of input context, according to benchmarking by Artificial Analysis. The device is designed to accelerate the decode phase of inference, the part users experience as latency while an AI agent thinks, calls tools, and streams responses.

Nebius Leads Deployment

Nebius will be the first AI cloud to bring Groq 3 LPX to production via its Token Factory platform. CTO Danila Shtan said the goal is to make every step of an agents loop feel instant through the same API developers already use, with no migration to a new stack. Following Nebius, the purpose-built inference cloud Groq plans to be among the earliest adopters.

The LPX extends the Vera Rubin NVL72 platform by adding LPU-based decode acceleration alongside GPUs for prefill, creating a heterogeneous architecture. NVIDIA says the combination of Rubin GPUs for prefill and Groq LPU for fast decode addresses the reality that one chip topology does not win both batch economics and interactive decode.

Implications for Agentic AI

The launch targets a specific pain point: interactive AI applications where decode latency directly determines user experience. Coding agents, live collaboration tools, and long-context applications that need real-time streaming benefit most from the accelerated token generation.

NVIDIA has not published per-token pricing for LPX. Access is expected through Nebius Token Factory tier selection, with premium pricing versus standard GPU tiers. Analysts note that the comparison should focus on cost per successful agent task rather than cost per million tokens, since faster decode may reduce user abandonment and improve task completion rates.

The launch sits in the broader context of NVIDIA expanding beyond hardware into the AI access layer, following its $20 billion acquisition of Groqs assets and its recent $6 billion commitment to Poolside for open-weight model development.

SourcesNVIDIA Newsroom; explainx.ai; StorageReview; Artificial Analysis
React to this dispatch
Share this dispatch X WhatsApp Bluesky Report an error
Written by

Founder and editor of Pulse of Nations, an independent wire service covering war, geopolitics, markets and technology.

discussion

Leave a Reply

Next dispatch Australia Bans AI-Generated Music From Official Charts Read →