Google released Gemma 4 12B, a multimodal model with an encoder-free architecture that processes vision and audio directly without separate encoding models, reducing latency. The model runs locally on standard devices with 16GB GPU memory and is the first medium-sized Gemma model capable of native audio input. Google is simultaneously launching desktop applications and LiteRT-LM local API server tools to enable developers to build offline agentic AI applications on consumer hardware.
Gemma 4 12B: The Developer Guide
calendar_today
June 3, 2026
person
Andre Susano Pinto, Andreas Steiner, Karolis Misiunas, Karsten Roth, Michael Tschannen, Omar Sanseviero
domain
gemini