Models

Liquid AI Debuts LFM2.5-VL-3B On-Device Vision Model

Liquid AI has launched LFM2.5-VL-3B, a 3.1-billion-parameter vision-language model designed to run locally on consumer devices while matching the performance of larger competitors.

MarkTechPost3 days agoModels
Image: MarkTechPost

Liquid AI has introduced LFM2.5-VL-3B, a 3.1-billion-parameter vision-language model optimized for on-device deployment. Built on the LFM2.5-2.6B language backbone and a SigLIP2 NaFlex shape-optimized 400-million-parameter vision encoder, the model is designed to read digital screens, ground objects to coordinates, and parse documents. It runs locally, fitting into roughly 3 gigabytes of memory. On an Apple M5 Max, it decodes at 228 tokens per second, while achieving 20 tokens per second on a Galaxy S26 Ultra.

The model achieves an average score of 69.4 across 28 vision benchmarks, matching the 4.7-billion-parameter InternVL-3.5-4B and trailing the Qwen3.5-4B by 0.7 points. On ScreenSpot-v2, LFM2.5-VL-3B averaged 80.7, comprising 78.7 on desktop, 81.2 on mobile, and 82.2 on web. This beats Gemma-4-E4B at 51.2 and Qwen3.5-4B at 78.5, though it trails InternVL-3.5-4B at 84.1. Other results include 73.1 on RealWorldQA, 84.3 on TextVQA, 63.3 on MMStar, 68.5 on MathVista-mini, 81.3 on ChartQA, 91.1 on DocVQA, and 84.2 on OCRBench v1. Its CountBenchQA score regressed from 92.2 to 87.3. On text-only evaluation, its IFEval score rose from 72.9 to 82.3, compared to Gemma-4-E4B at 87.9.

This release introduces function calling to the vision-language line, with ToolSandbox performance rising from 26.4 to 59.5 and BFCL v4 climbing from 20.5 to 32.5. Grounding precision on RefCOCO-avg jumped from 57.1 to 87.9. Multi-image capabilities also improved, with BLINK rising from 50.2 to 61.5 and MuirBench from 34.9 to 58.3. The model was pre-trained on approximately 34 trillion tokens, doubling its vocabulary to 128,000 tokens to support 16 languages. Post-training involved supervised fine-tuning with knowledge distillation and specialized training to prevent failures, followed by multi-reward reinforcement learning.

To support immediate deployment, Liquid AI is shipping the model in native, GGUF, ONNX, and MLX formats, compatible with runtimes like llama.cpp, vLLM, and SGLang. It is released under the LFM Open License v1.0, which allows free commercial use for companies with under 10 million dollars in annual revenue, while larger enterprises must negotiate custom terms.

This is our own summary of reporting by MarkTechPost

More in Models