NVIDIA Accelerates Google DeepMind’s DiffusionGemma for Local AI

Google DeepMind released DiffusionGemma, an experimental open model that denoises up to 256 tokens at once, and NVIDIA optimized it to run up to 4x faster on RTX and DGX systems.

Wednesday June 10, 2026 Source: nvidia.com
TL;DR — Quick Answer

Google DeepMind released DiffusionGemma on June 10, 2026, an experimental open model that generates text by denoising whole blocks of up to 256 tokens in parallel instead of one token at a time. It is built on Gemma 4, a 26-billion-parameter mixture-of-experts model that activates about 3.8 billion parameters per step, and is licensed under Apache 2.0. NVIDIA optimized it to run up to four times faster on its hardware, reaching about 1,000 tokens per second on an H100, 150 tokens per second on DGX Spark and up to 2,000 tokens per second on DGX Station.

Key Takeaways

NVIDIA Accelerates Google DeepMind’s DiffusionGemma for Local AI — AI news article illustration

Google DeepMind released DiffusionGemma on June 10, 2026, an experimental open model built for exceptionally fast text generation. NVIDIA immediately optimized it across GeForce RTX GPUs, the RTX PRO platform and DGX Spark systems, from local PCs to the cloud.

A Different Way to Generate Text

Almost every widely used large language model is autoregressive, emitting one token at a time and waiting for each before computing the next. DiffusionGemma borrows from image generation instead, starting from noise and refining a whole block at once, denoising up to 256 tokens per step.

Built on Gemma 4

The model sits on Gemma 4, a 26-billion-parameter mixture-of-experts architecture that activates only about 3.8 billion parameters per step, pairing a diffusion head with Google’s Gemma 4 design. It is released with open weights under a permissive Apache 2.0 license.

Performance on NVIDIA Hardware

Sequential decoding is memory-bound, leaving compute idle. Pulling a 256-token block through the transformer in parallel is compute-bound, which plays directly to GPU strengths.

Where to Run It

The model runs out of the box on a GeForce RTX 5090 or DGX Spark through Hugging Face Transformers, with day-zero serving in vLLM. Fine-tuning is available through Unsloth and the NVIDIA NeMo framework, and llama.cpp support is coming soon.

Why It Matters

Interactive chat, agentic loops and on-device assistants all stall when generation is sequential. A model that thinks in blocks changes that tradeoff, and because it runs locally there is no cloud round trip or per-token cost.

Frequently Asked Questions

What is DiffusionGemma?

DiffusionGemma is an experimental open text model from Google DeepMind, released June 10, 2026, that uses diffusion to denoise blocks of up to 256 tokens in parallel rather than predicting one token at a time.

How much faster is DiffusionGemma on NVIDIA hardware?

NVIDIA reports up to 4x faster single-user performance, with about 1,000 tokens per second on an H100 GPU, 150 tokens per second on DGX Spark and up to 2,000 tokens per second on DGX Station.

Is DiffusionGemma open source?

Yes, it is released with open weights under a permissive Apache 2.0 license, and it runs locally without cloud access or per-token costs.

Where can developers run DiffusionGemma?

It runs on GeForce RTX GPUs, the RTX PRO platform and DGX Spark systems, with support in Hugging Face Transformers, vLLM and Unsloth, plus NVIDIA-hosted APIs on build.nvidia.com.

This article is based on the official announcement from nvidia.com . Read the original for full technical details.

Related Articles

Back to all news