Gemma is Google DeepMind's family of lightweight, open AI models built from Gemini research and technology. The family includes general and specialized variants; current Gemma 4 variants are multimodal, supporting text and image input across the family and audio/video capabilities on supported models, with text output.
Pricing
Model Intelligence
Recent stories
Google’s new Gemma 4 12B ships as an encoder-free open model for text, image, audio, and video tasks with a 256K context window. Early GGUF ports and local benchmarks make it a plausible on-device multimodal option for creator tooling and experimentation.
Google DeepMind shipped four Gemma 4 models with multimodal input, including 31B Dense, 26B MoE, and two edge variants available through AI Studio, Hugging Face, Kaggle, and Ollama. Early community tests say local performance and usable context windows still vary by runtime, quantization, and GPU memory.