🔍 Read the full analysis: A Practical Look At Accelerating Vision-Language Models With LFM2.5-VL-DSpark on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Liquid AI released LFM2.5-VL-DSpark, an experimental 280M-parameter draft model for speeding up its LFM2.5-VL-3B vision-language model. The company reports decoding gains of up to 3.13x on Apple silicon and 2.66x on an H100, while end-to-end gains are lower; the measurements have not been independently verified.
Liquid AI released LFM2.5-VL-DSpark, an experimental draft model designed to speed up inference for its LFM2.5-VL-3B vision-language model. The company reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an NVIDIA H100, while saying the method preserves the target model’s output under greedy decoding.
The add-on contains about 279.5 million parameters, roughly 8.9% of the target model’s parameter count. It proposes short blocks of candidate tokens using hidden states from selected layers of the vision-language model. The target model then checks those candidates, accepting matching tokens and discarding the rest, a process related to practical AI decision models. This speculative-decoding process can reduce sequential token generation while retaining the target model’s greedy output.
Liquid AI reports results across six vision-language task categories, including general and text-based visual question answering, captioning, chart questions, complex reasoning, and multi-turn conversation. On an M5 Max using MLX, the company measured decoding gains of 2.30x to 3.13x and end-to-end latency gains of 1.56x to 2.62x. On an M3 Ultra using llama.cpp, it reported 1.57x to 2.14x faster decoding and 1.30x to 1.77x faster end-to-end performance.
For an H100 GPU, Liquid AI reported decoding improvements reaching 2.66x and end-to-end gains of 1.64x to 2.27x. The GPU results use a DSpark block size of eight, while the drafter’s described training configuration uses a block size of nine. Liquid AI says the model is available on Hugging Face in Safetensors and GGUF formats. It also points to support in llama.cpp, MLX-VLM, and SGLang, with particular integration changes required in each project.
Faster Local Vision-Language Responses
Vision-language models can spend substantial time processing an image and its associated visual tokens before generating a response. DSpark targets the token-generation stage, which can matter for interactive uses such as asking questions about a document, chart, or photograph. A smaller draft model that speeds decoding could make responses feel faster on supported computers without requiring a larger target model.
The reported end-to-end results are more modest than the decoding figures because DSpark does not accelerate every part of inference. Image encoding and prompt processing remain, so they continue to contribute to total latency. The difference between decode and end-to-end gains is relevant to users deciding whether the method improves the whole task, rather than only the model’s token-generation phase.
Availability in established inference projects may also lower the effort needed to try the approach. However, performance depends on the hardware, task, prompt, and image resolution. The published numbers are Liquid AI’s measurements; independent results would help establish how consistently the gains carry over to other setups.
DSpark Moves Beyond Text Models
Speculative decoding uses a smaller, faster model to propose tokens and a larger target model to verify them. When proposals match, the target can accept several tokens in a verification step. Under greedy decoding, this process is designed to preserve the target model’s output while changing how quickly tokens are produced.
Liquid AI says it previously applied its DSpark recipe to text-only LFM2.5 models. This release adapts the method to a multimodal target. According to the company, image patches and text tokens enter a shared representation before the layers whose hidden states the drafter uses. That lets the drafter work with vectors of the same dimensionality across both modalities, using an inference algorithm the company says is unchanged from its text models.
The released drafter is described as an attention-only model with four layers, selected after comparisons involving three, four, and five layers. Liquid AI says it was trained on a mixture of vision-language supervised fine-tuning data weighted toward expected serving workloads. The company has not published the dataset mixture in detail.
“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”
— Liquid AI
Benchmark Scope and Open Questions
The release is labeled experimental, and Liquid AI has not said when or whether it will become a stable release. The performance figures are company-reported rather than independently verified, and actual results may vary with hardware, prompt composition, image resolution, and task.
One GPU benchmark range in the source material gives a lower bound of “20.4x” and an upper bound of 2.66x. The lower figure appears inconsistent with the stated range, but Liquid AI has not clarified it; it should not be treated as a reliable result. The company also has not provided runtime memory measurements beyond the drafter’s parameter count.
Further questions include how well token proposals hold up on vision tasks outside the training mixture, and how the method behaves with sampling-based generation rather than greedy decoding. The stated exactness applies to greedy output; the source does not establish equivalent output behavior for other decoding settings. The training-data composition and independent benchmark results also remain unavailable.
Independent Tests and Release Status
The next useful evidence will come from independent benchmarks across hardware and representative vision-language tasks, including measurements of total latency and runtime memory. Such tests could clarify how much of the reported decoding gain users see in complete requests.
Adoption will also depend on the state of the integrations. Liquid AI identifies changes for llama.cpp, MLX-VLM, and SGLang, with SGLang requiring a build that supports DSpark for LFM2 targets. The company has not announced a schedule for a stable release or further benchmark clarification.
Key Questions
What is LFM2.5-VL-DSpark?
It is an experimental draft model that proposes tokens for Liquid AI’s LFM2.5-VL-3B vision-language model to verify as part of speculative decoding.
Does DSpark change the model’s answers?
Liquid AI says greedy decoding produces the same output as the target model alone because the target verifies proposed tokens. The source material does not establish the same behavior for sampling-based generation.
How much faster is it?
Liquid AI reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an H100. End-to-end improvements are lower in its reported tests, and results have not been independently verified.
Where can developers try it?
Liquid AI says the model is available on Hugging Face in Safetensors and GGUF formats, with integration support identified for llama.cpp, MLX-VLM, and SGLang.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
