Groq Among First to Deploy NVIDIA Groq 3 LPX and Vera Rubin NVL72 for Inference

Groq announces it will be among the first adopters of NVIDIA Groq 3 LPX with Vera Rubin NVL72, deploying through Dell Technologies to power its inference cloud with 3,400 tokens/sec on agentic workloads.

Monday August 24, 2026 Source: Groq
TL;DR — Quick Answer

Groq announced it will be among the first to deploy NVIDIA Groq 3 LPX with Vera Rubin NVL72 — a striking pivot for the company that built its brand on LPUs as the NVIDIA alternative. The platform hits 3,400 tokens/sec on Gemma 4 31B in benchmarks, deployed with Dell, following Groq's $350M Series A and $650M raise in June.

Key Takeaways

Groq Among First to Deploy NVIDIA Groq 3 LPX and Vera Rubin NVL72 for Inference — AI news article illustration

Groq has announced it will be among the first adopters of NVIDIA Groq 3 LPX, boosting inference token generation for NVIDIA Vera Rubin NVL72 systems deployed to its purpose-built AI inference cloud. Groq is working with Dell Technologies to deploy the new platform.

The Performance Numbers

NVIDIA’s published benchmarks for the Vera Rubin platform with Groq 3 LPX include:

A Notable Pivot

The announcement is striking given Groq’s history: the company built its brand on its own LPU (Language Processing Unit) chips as an alternative to NVIDIA GPUs. Now Groq is deploying NVIDIA’s newest inference platform alongside its LPU fleet — a pragmatic move that reflects how the inference market is evolving.

Groq notes it is “the only team with hands-on experience operating LPUs in production at scale,” positioning that expertise as complementary to the NVIDIA deployment. The company became an NVIDIA Cloud Partner in August, committing to NVIDIA’s reference architecture and operational standards.

Scale Context

Groq’s inference cloud serves:

The company also recently closed a $350 million Series A (August 17) to scale its AI inference cloud business, following a $650 million raise in June.

What It Means

For Groq customers, the NVIDIA deployment adds capacity on infrastructure “already optimized for high-demand inference workloads.” For the industry, it signals that the inference market is big enough — and competitive enough — that even dedicated chip startups are hedging their infrastructure bets. Speed, it turns out, is the product — regardless of whose silicon delivers it.

Frequently Asked Questions

What is NVIDIA Groq 3 LPX?

NVIDIA Groq 3 LPX is an inference acceleration platform that dramatically increases token generation rates for NVIDIA Vera Rubin NVL72 systems — reaching 3,400 output tokens per second on Gemma 4 31B with 100K token context in Artificial Analysis benchmarking.

Why is Groq using NVIDIA chips?

Despite building its brand on custom LPU processors, Groq is deploying NVIDIA's newest inference platform pragmatically — reflecting how the inference market demands infrastructure diversity. Groq notes it remains the only team with LPU production experience at scale, positioning that expertise as complementary.

How big is Groq's inference cloud?

GroqCloud serves more than 6 million developers, Fortune 500 enterprises, and thousands of AI-native companies, generating trillions of tokens weekly across data centers in North America, Europe, the Middle East, and Asia-Pacific.

This article is based on the official announcement from Groq . Read the original for full technical details.

Related Articles

Back to all news