Gemma 3n E2B
Google·
LLMsopen-weightMobile / NPU
Overview
Phone-first open-weight model built on the MatFormer (Matryoshka Transformer) architecture, where the larger E4B model contains a fully functional E2B sub-model trained alongside it. Per-layer embeddings keep the resident footprint near the 2B effective parameters rather than the 5B raw ones, so it runs in about 2GB of memory on a handset. Handles text, image and audio input.
Capabilities and innovations
Runs in ~2GB memory on a phoneText, image and audio inputMatFormer elastic inferenceFully offline via LiteRT / Google AI EdgeMatFormer nested sub-model (E2B inside E4B)Per-layer embeddings cut resident memory
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| MMLU | 60.1% | vendor 60.1% MMLU (0-shot, instruction-tuned) for the E2B variant. |
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
Architecture and hardware
- Parameters
- 5B total, 2B active per token (MatFormer)
- Estimated VRAM at Q4
- ~3 GB, Mobile / NPU class
- Quantization formats
- GGUF, LiteRT, INT4
- Recommended runtime
- Ollama
- License
- Gemma Terms of Use
Links
More from Google
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.