AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemma 4 12B

Google·

LLMsopen-weightEdge / Consumer

Overview

Open-weight 12B multimodal model from Google, released under Apache 2.0. Processes text, images, audio and video in a single encoder-free transformer with a 256K-token context window, and runs locally on a machine with 16GB of RAM or VRAM. On 2026-06-05 Google released Quantization-Aware Training (QAT) checkpoints across the Gemma 4 family (Q4_0 plus a mobile-specialized format), which Google reports cuts memory use by roughly 72%, letting the 12B run on consumer GPUs with little quality loss.

Capabilities and innovations

Encoder-free unified multimodal (text, image, audio, video)256K-token context windowRuns on 16GB RAM/VRAMApache 2.0 licenseSingle encoder-free multimodal transformerQuantization-Aware Training checkpoints (~72% less memory)Local multimodality at the 12B size

Benchmarks

No benchmark values are recorded for this model.

Architecture and hardware

Parameters
12B total (Dense)
Estimated VRAM at Q4
~7 GB, Edge / Consumer class
Quantization formats
GGUF, QAT Q4_0, BF16
Recommended runtime
Ollama
License
Apache 2.0

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.