DiffusionGemma
An experimental open diffusion model for faster block-based generation.
An experimental open, multimodal diffusion-model family in the Gemma line that generates and refines blocks of text rather than decoding strictly one token at a time; the announced 26B-A4B member is positioned for faster generation and long-context, image, and video use.
Pricing
Model Intelligence
Recent stories
Google released Apache 2.0 DiffusionGemma, a 26B-A4B diffusion text model that claims up to 4x faster output by generating text in blocks instead of one token at a time. The release matters for local and hosted stacks that want to test a new decoding path.
Google's new diffusion text model picked up same-day runtime support: vLLM added native diffusion-LM serving, Unsloth shipped GGUFs, and llama.cpp got local setup guidance. That shortens the path from release to local and hosted evaluation.