A Practical Look At Accelerating Vision-Language Models With LFM2.5-VL-DSpark
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Practical Look At Accelerating Vision-Language Models With LFM2.5-VL-DSpark on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Liquid AI released LFM2.5-VL-DSpark, an experimental 280M-parameter draft model for speeding up its LFM2.5-VL-3B vision-language model. The company reports decoding gains of up to 3.13x on Apple silicon and 2.66x on an H100, while end-to-end gains are lower; the measurements have not been independently verified.

Liquid AI released LFM2.5-VL-DSpark, an experimental draft model designed to speed up inference for its LFM2.5-VL-3B vision-language model. The company reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an NVIDIA H100, while saying the method preserves the target model’s output under greedy decoding.

The add-on contains about 279.5 million parameters, roughly 8.9% of the target model’s parameter count. It proposes short blocks of candidate tokens using hidden states from selected layers of the vision-language model. The target model then checks those candidates, accepting matching tokens and discarding the rest, a process related to practical AI decision models. This speculative-decoding process can reduce sequential token generation while retaining the target model’s greedy output.

Liquid AI reports results across six vision-language task categories, including general and text-based visual question answering, captioning, chart questions, complex reasoning, and multi-turn conversation. On an M5 Max using MLX, the company measured decoding gains of 2.30x to 3.13x and end-to-end latency gains of 1.56x to 2.62x. On an M3 Ultra using llama.cpp, it reported 1.57x to 2.14x faster decoding and 1.30x to 1.77x faster end-to-end performance.

For an H100 GPU, Liquid AI reported decoding improvements reaching 2.66x and end-to-end gains of 1.64x to 2.27x. The GPU results use a DSpark block size of eight, while the drafter’s described training configuration uses a block size of nine. Liquid AI says the model is available on Hugging Face in Safetensors and GGUF formats. It also points to support in llama.cpp, MLX-VLM, and SGLang, with particular integration changes required in each project.

At a glance
announcementWhen: Released in 2026; available now, with e…
The developmentLiquid AI has extended its DSpark speculative-decoding approach to its open-weight LFM2.5-VL-3B vision-language model.
At a glance
announcementWhen: announced September 2026, available now
The developmentLiquid AI announced and released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, available immediately on Hugging Face.

Faster Local Vision-Language Responses

Vision-language models can spend substantial time processing an image and its associated visual tokens before generating a response. DSpark targets the token-generation stage, which can matter for interactive uses such as asking questions about a document, chart, or photograph. A smaller draft model that speeds decoding could make responses feel faster on supported computers without requiring a larger target model.

The reported end-to-end results are more modest than the decoding figures because DSpark does not accelerate every part of inference. Image encoding and prompt processing remain, so they continue to contribute to total latency. The difference between decode and end-to-end gains is relevant to users deciding whether the method improves the whole task, rather than only the model’s token-generation phase.

Availability in established inference projects may also lower the effort needed to try the approach. However, performance depends on the hardware, task, prompt, and image resolution. The published numbers are Liquid AI’s measurements; independent results would help establish how consistently the gains carry over to other setups.

DSpark Moves Beyond Text Models

Speculative decoding uses a smaller, faster model to propose tokens and a larger target model to verify them. When proposals match, the target can accept several tokens in a verification step. Under greedy decoding, this process is designed to preserve the target model’s output while changing how quickly tokens are produced.

Liquid AI says it previously applied its DSpark recipe to text-only LFM2.5 models. This release adapts the method to a multimodal target. According to the company, image patches and text tokens enter a shared representation before the layers whose hidden states the drafter uses. That lets the drafter work with vectors of the same dimensionality across both modalities, using an inference algorithm the company says is unchanged from its text models.

The released drafter is described as an attention-only model with four layers, selected after comparisons involving three, four, and five layers. Liquid AI says it was trained on a mixture of vision-language supervised fine-tuning data weighted toward expected serving workloads. The company has not published the dataset mixture in detail.

“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”

— Liquid AI

Benchmark Scope and Open Questions

The release is labeled experimental, and Liquid AI has not said when or whether it will become a stable release. The performance figures are company-reported rather than independently verified, and actual results may vary with hardware, prompt composition, image resolution, and task.

One GPU benchmark range in the source material gives a lower bound of “20.4x” and an upper bound of 2.66x. The lower figure appears inconsistent with the stated range, but Liquid AI has not clarified it; it should not be treated as a reliable result. The company also has not provided runtime memory measurements beyond the drafter’s parameter count.

Further questions include how well token proposals hold up on vision tasks outside the training mixture, and how the method behaves with sampling-based generation rather than greedy decoding. The stated exactness applies to greedy output; the source does not establish equivalent output behavior for other decoding settings. The training-data composition and independent benchmark results also remain unavailable.

Independent Tests and Release Status

The next useful evidence will come from independent benchmarks across hardware and representative vision-language tasks, including measurements of total latency and runtime memory. Such tests could clarify how much of the reported decoding gain users see in complete requests.

Adoption will also depend on the state of the integrations. Liquid AI identifies changes for llama.cpp, MLX-VLM, and SGLang, with SGLang requiring a build that supports DSpark for LFM2 targets. The company has not announced a schedule for a stable release or further benchmark clarification.

Key Questions

What is LFM2.5-VL-DSpark?

It is an experimental draft model that proposes tokens for Liquid AI’s LFM2.5-VL-3B vision-language model to verify as part of speculative decoding.

Does DSpark change the model’s answers?

Liquid AI says greedy decoding produces the same output as the target model alone because the target verifies proposed tokens. The source material does not establish the same behavior for sampling-based generation.

How much faster is it?

Liquid AI reports decoding speedups of up to 3.13x on an M5 Max and up to 2.66x on an H100. End-to-end improvements are lower in its reported tests, and results have not been independently verified.

Where can developers try it?

Liquid AI says the model is available on Hugging Face in Safetensors and GGUF formats, with integration support identified for llama.cpp, MLX-VLM, and SGLang.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Signal: Four Frontier-Class Open Models in Eight Weeks — China’s Release Cadence Is the Story

Chinese AI labs launched four frontier-class open models within eight weeks, signaling a rapid production line that challenges Western dominance.

The Free-Download Question: When Running Your Own Model Actually Beats Paying

Analysis of when owning and running open-weight models becomes more cost-effective than paying for API access, considering hardware, operation, and performance.

How IBM Time Series Models Enable Instant AI Insights On Confluent Platform

IBM Granite Time Series models now available in Early Access on Confluent Cloud, enabling real-time forecasting and anomaly detection within Apache Flink.

The Question No To-Do App Can Answer

Thorsten Meyer AI says Threlmark ranks work across projects, adds flow signals, and supports AI agent handoffs.