Back to Articles

LFM2.5-VL-3B is our most capable vision-language model you can run on your own hardware. It understands documents and screens alike, grounds objects, and can call tools. It answers directly instead of reasoning, so responses stay fast in real-time and on-device apps.

LFM2.5-VL-3B extends the vision-language capabilities of our previous releases with four major improvements:

Screen/UI understanding: Strong understanding of digital screens across different devices.

Grounding: Improved grounding and object detection with natural language queries.