Liquid AI released LFM2.5-2.6B, an agentic model that runs entirely on-device. It plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots. The model has 2.69B total parameters, a 131,072-token context window, and a 128,000-token vocabulary. Pre-training used approximately 34 trillion tokens. Two checkpoints shipped: LFM2.5-2.6B-Base for fine-tuning, and LFM2.5-2.6B post-trained for agentic workloads. Because inference stays local, data never leaves the device and the marginal cost of each run is near zero. Liquid AI reports tool-use and instruction-following scores competitive with models nearly four times its size.
Is it deployable
The answer is Yes. Both checkpoints are public on Hugging Face under the lfm1.0 license. Weights ship in native, GGUF, MLX, and ONNX formats, with day-one support in llama.cpp, vLLM, SGLang, and LM Studio.
Which companies: Solo developers and startups can pilot on hardware they already own. The model decodes at 220 tokens/s on an M5 Max in under 2.5 GB. Mid-market teams can self-host on one GPU: a single NVIDIA H100 SXM5 serves roughly 1.3B tokens per day. Enterprises and OEMs can push the same weights to device fleets through GGUF and ONNX. Fine-tuning is available via LoRA with TRL and Unsloth.










