Google DeepMind released DiffusionGemma on June 10, 2026, an experimental open model built for exceptionally fast text generation. NVIDIA immediately optimized it across GeForce RTX GPUs, the RTX PRO platform and DGX Spark systems, from local PCs to the cloud.
A Different Way to Generate Text
Almost every widely used large language model is autoregressive, emitting one token at a time and waiting for each before computing the next. DiffusionGemma borrows from image generation instead, starting from noise and refining a whole block at once, denoising up to 256 tokens per step.
Built on Gemma 4
The model sits on Gemma 4, a 26-billion-parameter mixture-of-experts architecture that activates only about 3.8 billion parameters per step, pairing a diffusion head with Google’s Gemma 4 design. It is released with open weights under a permissive Apache 2.0 license.
Performance on NVIDIA Hardware
Sequential decoding is memory-bound, leaving compute idle. Pulling a 256-token block through the transformer in parallel is compute-bound, which plays directly to GPU strengths.
- Around 1,000 tokens/sec on a single NVIDIA H100
- 150 tokens/sec on NVIDIA DGX Spark
- Up to 2,000 tokens/sec on NVIDIA DGX Station
- Roughly 4x faster than an equivalent autoregressive model in the same single-user regime
Where to Run It
The model runs out of the box on a GeForce RTX 5090 or DGX Spark through Hugging Face Transformers, with day-zero serving in vLLM. Fine-tuning is available through Unsloth and the NVIDIA NeMo framework, and llama.cpp support is coming soon.
Why It Matters
Interactive chat, agentic loops and on-device assistants all stall when generation is sequential. A model that thinks in blocks changes that tradeoff, and because it runs locally there is no cloud round trip or per-token cost.