AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Gemma 3n E2B

Google·

LLMsopen-weightMobile / NPU

Overview

Phone-first open-weight model built on the MatFormer (Matryoshka Transformer) architecture, where the larger E4B model contains a fully functional E2B sub-model trained alongside it. Per-layer embeddings keep the resident footprint near the 2B effective parameters rather than the 5B raw ones, so it runs in about 2GB of memory on a handset. Handles text, image and audio input.

Capabilities and innovations

Runs in ~2GB memory on a phoneText, image and audio inputMatFormer elastic inferenceFully offline via LiteRT / Google AI EdgeMatFormer nested sub-model (E2B inside E4B)Per-layer embeddings cut resident memory

Benchmarks

BenchmarkScoreSource
MMLU60.1%vendor
60.1% MMLU (0-shot, instruction-tuned) for the E2B variant.
  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor

Architecture and hardware

Parameters
5B total, 2B active per token (MatFormer)
Estimated VRAM at Q4
~3 GB, Mobile / NPU class
Quantization formats
GGUF, LiteRT, INT4
Recommended runtime
Ollama
License
Gemma Terms of Use

Links

More from Google

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.