Artificial Intelligence
August 24, 2026
Aug. 24, 2026 — NVIDIA today announced that NVIDIA Groq 3 LPX, the interactive AI inference accelerator, is now in full production. An extension of the NVIDIA Vera Rubin platform, Groq 3 LPX delivers a major boost in AI inference by enabling ultrafast token generation for highly responsive agentic systems.
Groq 3 LPX extends the Vera Rubin Platform, delivering ultrafast token generation
Agentic systems can generate massive volumes of tokens across hundreds or thousands of inference steps, making faster token generation critical for agents to reason, act and complete complex tasks in real time.
Vera Rubin NVL72 systems provide the most versatile training and inference platform for every AI factory. NVIDIA Groq 3 LPX extends the inference performance of Vera Rubin NVL72 by dramatically increasing the rate of token generation, providing premium user experiences for context-heavy workloads so agents can act at extreme speeds.
NVIDIA Groq 3 LPX is pushing the frontier of AI inference. It delivered a record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B, an openntic systems — the fastest performance ever recorded for the model
Groq 3 LPX enables agentic tasks such as coding in minutes versus hours, providing 4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform.
“Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” said Jensen Huang, founder and CEO of NVIDIA. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide.”
Groq 3 LPX — The Interactive AI Inference Accelerator
Agentic AI creates two distinct computing challenges: efficiently processing enormous amounts of context and generating tokens with extremely low latency.
NVIDIA Groq 3 LPX is purpose-built to extend Vera Rubin’s interactivity — the rate at which tokens are generated for an individual user, determining how quickly an agent can complete each step of its work.
Faster generation gives agents more time to inspect files, write and test code, call tools, verify results and iterate while maintaining a responsive user experience.
AI Cloud Momentum for Groq 3 LPX
AI clouds are becoming the engines of the AI economy, giving enterprises and developers access to advanced infrastructure for training, reasoning and inference at scale. For providers serving latency-sensitive, high-volume inference workloads, NVIDIA Groq 3 LPX provides a path to deploy differentiated compute in proven rack-scale systems.
Nebius, a leading AI cloud, plans to bring NVIDIA Groq 3 LPX to Nebius Token Factory, its production inference platform, giving developers access to extreme token generation speed for highly responsive agentic AI applications.
“Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Danila Shtan, chief technology officer of Nebius. “As the first AI cloud bringing it to productionnt’s loop feels instant — through the same API developers are already using, with no migration to a new stack.”
Following Nebius, purpose-built AI inference cloud Groq plans to be among the platform’s earliest adopters.
Extreme Codesign for AI Factories
Through extreme codesign across seven chips and five purpose-built racks, NVIDIA Vera Rubin is the most extensive AI factory platform.
NVIDIA Vera Rubin NVL72 and Groq 3 LPX tackle the various workload requirements of customer AI factories, including frontier model makers and open model service providers.
These rack platforms feature NVIDIA BlueField-4 DPUs and work in combination with NVIDIA Vera CPU racks, NVIDIA Vera BlueField-4 STX storage and NVIDIA Spectrum-6 SPX Ethernet to optimize multi-agent systems for the highest throughput per watt and the lowest-latency inference.
NVIDIA (NASDAQ: NVDA) is the world leader in AI and accelerated computing.
The Big Data Inside Amazon’s New Fire Phone
The new Fire phone that Amazon launched this week looks like your ordinary black smartphone,…
Three Reasons to be Scared of the Internet of Things
We know the Internet of Things forecasts: 50 billion connected devices by 2020. Apparently, there’s…
Wanted: Intelligent Middleware That Simplifies Big Data Analytics
We’ve seen tremendous technological innovation in the data analytics space over the past 10 years….
GPUs Tackle Massive Data of the Hive Mind
LIVE from GTC12 — The flock of birds that weaves seamlessly through the sky, propelled…
Why Hadoop on IBM Power
In the quest to achieve data-driven insight, Hadoop running on Intel X86-based processors has emerged…
Datanami’s Leverage Big Data Summit Wraps Up
Dialog and networking were on the Datanami agenda this week as we kicked off our…
Why Enterprise AI Costs Are an Inference Problem, Not a Training One
Ask most enterprises where their AI budget goes and they point at training runs, or…
Data Is the New Gold? YipitData Aims for $2.5B-Plus Sale
The AI boom has mostly been a race for more chips, better models and bigger…
Nobel Laureate David Baker Takes Aim at the Virtual Cell with GenBio AI
What if scientists could test a new drug or make changes to a cell without…
Oakley Capital Bets Big on Graphwise to Solve a Growing AI Problem
European private equity firm Oakley Capital has acquired a majority stake in Graphwise, a knowledge…
AI Agents Are Creating a New Data Problem. Ciklum and ClickHouse Have a Plan
Technology services company Ciklum has formed a strategic partnership with real-time analytics database provider ClickHouse,…
AI Is Forcing Analytics Teams Into a New Role
For years, analytics teams have been responsible for helping organizations understand their data. They built…
