Groq has announced it will be among the first adopters of NVIDIA Groq 3 LPX, boosting inference token generation for NVIDIA Vera Rubin NVL72 systems deployed to its purpose-built AI inference cloud. Groq is working with Dell Technologies to deploy the new platform.
The Performance Numbers
NVIDIA’s published benchmarks for the Vera Rubin platform with Groq 3 LPX include:
- 3,400 output tokens per second running Gemma 4 31B with 100K token context in Artificial Analysis benchmarking
- Agentic tasks like coding in minutes versus hours
- 4x higher interactivity for latency-sensitive agentic AI workloads than the nearest alternative platform
A Notable Pivot
The announcement is striking given Groq’s history: the company built its brand on its own LPU (Language Processing Unit) chips as an alternative to NVIDIA GPUs. Now Groq is deploying NVIDIA’s newest inference platform alongside its LPU fleet — a pragmatic move that reflects how the inference market is evolving.
Groq notes it is “the only team with hands-on experience operating LPUs in production at scale,” positioning that expertise as complementary to the NVIDIA deployment. The company became an NVIDIA Cloud Partner in August, committing to NVIDIA’s reference architecture and operational standards.
Scale Context
Groq’s inference cloud serves:
- More than 6 million developers
- Fortune 500 enterprises and thousands of AI-native companies
- Trillions of tokens generated every week across North America, Europe, the Middle East, and APAC
The company also recently closed a $350 million Series A (August 17) to scale its AI inference cloud business, following a $650 million raise in June.
What It Means
For Groq customers, the NVIDIA deployment adds capacity on infrastructure “already optimized for high-demand inference workloads.” For the industry, it signals that the inference market is big enough — and competitive enough — that even dedicated chip startups are hedging their infrastructure bets. Speed, it turns out, is the product — regardless of whose silicon delivers it.