Google Deepmind's Gemma 4 12B is an open-source model that processes text, images, and audio natively and runs on laptops with just 16 GB of RAM. It nearly matches the twice-as-large 26B model in benchmarks and ships under an Apache 2.0 license for commercial use.

Meet Gemma 4 12B: the first medium-sized, encoder-free multimodal model capable of natively ingesting audio and video. Ideal for local AI development with 16GB VRAM, Hugging Face…

An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.

An in-depth explainer to Gemma 4 12B; a unified, encoder-free multimodal model!