AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

LFM2.5-2.6B

Liquid AI·

LLMsopen-weightMobile / NPU

Overview

On-device agentic model from Liquid AI, 2.69B parameters with a 131K-token context window and tool calling, trained to work inside real agent harnesses rather than to top knowledge benchmarks. Liquid measures 220 tok/s on an Apple M5 Max, 113 tok/s on an AMD Ryzen AI Max+ 395 and roughly 30 tok/s on a phone, all under 2.5GB of memory, and reports it beating the 4x larger Qwen3.5-9B on ToolSandbox (77.83 vs 76.44) and on every instruction-following benchmark. No MMLU or GPQA figures have been published, so the benchmark fields stay null.

Capabilities and innovations

Runs under 2.5GB memory, ~30 tok/s on a phone131K-token context windowTool calling for on-device agentsOpen weightsAgent-harness training at 2.6B scaleBeats a 4x larger model on tool-use benchmarks

Benchmarks

No benchmark values are recorded for this model.

  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor

Architecture and hardware

Parameters
2.69B total (Hybrid)
Estimated VRAM at Q4
~2 GB, Mobile / NPU class
Quantization formats
GGUF, INT4
Recommended runtime
llama.cpp
License
LFM Open License v1.0

Links

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.