Back to Articles
LFM2.5-2.6B is built to power capable agents entirely on-device. It supports tool calling and multi-step workflows while staying small and fast enough for everyday hardware, from laptops to phones. This enables developers to deploy agents everywhere, keep data private on the device, and scale usage without a cloud inference bill.
Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks.
Agentic reinforcement learning: Trained inside the most popular agentic harnesses to improve compatibility.
Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory.







