Gemma 4, a multimodal model from Google DeepMind, is now in private preview on Cerebras Inference with general availability later this month. The model runs at over 1,500 tokens per second, roughly 15x faster than comparable production alternatives, enabling image-based tasks like screenshot analysis and document understanding for computer vision and agentic workflows. It is open-weight under the Apache 2.0 license and positioned as a reference medium-sized model on the platform.