Gemma 4 12B
Google·
LLMsopen-weightEdge / Consumer
Overview
Open-weight 12B multimodal model from Google, released under Apache 2.0. Processes text, images, audio and video in a single encoder-free transformer with a 256K-token context window, and runs locally on a machine with 16GB of RAM or VRAM. On 2026-06-05 Google released Quantization-Aware Training (QAT) checkpoints across the Gemma 4 family (Q4_0 plus a mobile-specialized format), which Google reports cuts memory use by roughly 72%, letting the 12B run on consumer GPUs with little quality loss.
Capabilities and innovations
Encoder-free unified multimodal (text, image, audio, video)256K-token context windowRuns on 16GB RAM/VRAMApache 2.0 licenseSingle encoder-free multimodal transformerQuantization-Aware Training checkpoints (~72% less memory)Local multimodality at the 12B size
Benchmarks
No benchmark values are recorded for this model.
Architecture and hardware
- Parameters
- 12B total (Dense)
- Estimated VRAM at Q4
- ~7 GB, Edge / Consumer class
- Quantization formats
- GGUF, QAT Q4_0, BF16
- Recommended runtime
- Ollama
- License
- Apache 2.0
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.