LiquidAI Releases LFM2.5-VL-DSpark Model with Up to 3.13x Speedup
- Published
- Sep 24, 2026 — 14:08 UTC
On September 24, 2026, LiquidAI launched the LFM2.5-VL-DSpark model, which features an increase of 280 million parameters, representing an 8.9% increase over the previous 3 billion target. The model achieves up to 3.13x decode speedup on device and up to 2.66x on NVIDIA H100 GPUs. End-to-end latency improvements range from 1.56x to 2.62x for the MLX software tool on the M5 Max device, while llama.cpp on the M3 Ultra shows speedups of 1.57x to 2.14x. LiquidAI noted that speculative decoding enhances only the decode process, not vision encoding or prefill. The LFM2.5-VL-DSpark model is available on Hugging Face in Safetensors and GGUF formats. This follows earlier announcements regarding LFM2.5's performance enhancements, including a previous article on August 20, 2026, discussing up to 3.2x faster inference capabilities.
By Callan Zhang · Sep 24, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: Hugging Face Blog
