Google has released DiffusionGemma, an experimental open-weight language model that generates entire 256-token blocks simultaneously rather than word by word, delivering speeds of more than 1,000 tokens per second on a single NVIDIA H100 and over 700 tokens per second on an RTX 5090.
The model, introduced by Google DeepMind, is a 26-billion-parameter mixture-of-experts architecture that activates only 3.8 billion parameters per forward pass. Quantized to lower precision, it fits within 18 GB of VRAM on high-end consumer GPUs. Rather than predicting one token at a time from left to right, DiffusionGemma starts with a canvas of 256 random placeholder tokens and refines them across multiple passes until readable text emerges, borrowing the diffusion approach from image generation. Google says the resulting throughput on a dedicated GPU runs roughly four times faster than an equivalent autoregressive Gemma 4 model in single-user inference.
The speed gain has a specific hardware basis. Autoregressive inference on a dedicated GPU is typically memory-bandwidth-bound: the processor spends most of its time waiting on data transfers rather than computing. DiffusionGemma shifts the bottleneck toward raw compute by giving the GPU a block of 256 tokens to process in parallel. Google notes this advantage is specific to local, low-concurrency inference; in high-volume cloud serving, autoregressive models already saturate hardware efficiently, and DiffusionGemma can raise serving costs. Apple Silicon users are also unlikely to see the same speedup because unified-memory architectures are already bandwidth-limited during inference, Google says.
The speed comes at a quality trade-off. Google's own benchmarks show DiffusionGemma runs roughly 3.5 times faster than same-size Gemma 4 but scores lower on accuracy tests. Google positions it as a tool for researchers and developers exploring speed-critical, interactive local workflows rather than a production replacement for standard Gemma 4 models. The bi-directional attention, where every token can reference every other token during generation including later ones, is the meaningful differentiator: it makes DiffusionGemma suited for tasks that do not work well left to right, such as code infilling, in-line text editing, filling gaps in structured data like amino acid sequences, and tasks where each answer depends on future answers. An Unsloth fine-tune solving Sudoku grids is Google's example.
The weights are available on Hugging Face under an Apache 2.0 license. NVIDIA optimized DiffusionGemma for Hopper and Blackwell server hardware including DGX Spark and DGX Station, and quantized it for the GeForce RTX 5090 and 4090 for consumer use. Supported inference tools at launch include Hugging Face Transformers, vLLM with Red Hat integration, and MLX. Google's JAX fine-tuning toolkit Hackable Diffusion and Unsloth are available for training; llama.cpp support is planned. DiffusionGemma is also accessible through the Gemini Enterprise Agent Platform Model Garden and NVIDIA NIM.













