Skip to content
HChuggingface.co·

Accelerating vision-language models with LFM2.5-VL-DSpark

AI summary

Hugging Face has released an experimental DSpark draft model for their LFM2.5-VL-3B vision-language model. This new model, LFM2.5-VL-DSpark, incorporates a speculative decoding path to accelerate performance. It achieves a significant speedup with only a minimal increase in memory footprint, while maintaining the original output quality. This advancement aims to accelerate vision-language models on edge devices and beyond.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 24, 2026, 14:08 UTC

IngestedOffset at this time: UTC+0Sep 24, 2026, 15:02 UTC

Published
Sep 24, 2026, 14:08
Ingested
Sep 24, 2026, 15:02
Source type
Official
Tier
First-party
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Today, we release an experimental DSpark draft model for our vision-language model (VLM) LFM2.5-VL-3B . As with our recently released LFM2.5-DSpark drafter models , it adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.

- Faster inference: decode speedups up to 3.13x on device and 2.66x on an H100, with end-to-end gains up to 2.62x and 2.27x.

- Small memory cost: the drafter adds 280M parameters, 8.9% on top of the 3B target

- Day-one support: LFM-compatible DSpark integrations for llama.cpp, MLX-VLM, and SGLang

How does speculative decoding work for VLMs

The vision drafter uses the same architecture as our text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to draft a block of k candidate tokens. Image patches and text tokens are projected into a shared representation before those layers, so the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality. The inference algorithm is therefore unchanged from the text models.

Training and Architecture

We follow the DSpark recipe with a mixture of vision-language SFT data, weighted toward the workloads we expect the model to serve. Based on ablations across 3, 4, and 5 layers, the draft model is a simplified attention-only drafter with 4 layers and a block size of 9. We ran 10 epochs on the final mixture and measured acceptance after each, which improved with additional training tokens before reaching diminishing returns. At inference time, we recommend a block size of 8 or 9 depending on the hardware.

The resulting drafter has approximately 280M parameters and increases the deployed model’s parameter count by just 8.9%.

Component LFM2.5-VL-3B

Decoder stack (4 layers) 193.0M

Hidden-state projection 21.0M

Markov head 65.5M

Norms + confidence head 6.4k

Total 279.5M

Inference Speedup on CPU and GPU

The DSpark draft model for LFM2.5-VL-3B ships with day-one support for llama.cpp , MLX-VLM , and SGLang .

We measure both on-device inference and GPU inference. Both configurations use a DSpark block size of 8 and are evaluated on six diverse vision-based tasks, including general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation, following the MMSpec benchmark .

On-device inference. With MLX on an M5 Max, decoding runs 2.30x to 3.13x faster by task. End-to-end latency improves by 1.56x to 2.62x. With llama.cpp on an M3 Ultra, decoding improves by 1.57x to 2.14x and end-to-end by 1.30x to 1.77x.

GPU inference. On H100, the same drafter delivers 20.4x to 2.66x faster decoding, with end-to-end improvements of 1.64x to 2.27x.

Limitations of speculation for vision workloads

In LLMs, prefill is mostly compute-bound, and its cost grows (sub)quadratically with prompt length. VLMs add to this because the image first passes through a vision encoder, then the language backbone processes hundreds of visual tokens along with the text prompt. Edge devices have far less compute than datacenter GPUs, so prefill takes up more of the end-to-end latency, as time-to-first-token and decode measurements on Apple silicon and H100 show. (The M5's per-core GPU neural accelerators narrow this gap).

Speculative decoding speeds up only decode, not vision encoding or prefill. When those stages already take up much of the wall time, even a large decode speedup gives only a modest end-to-end gain. This is Amdahl's law, where the overall speedup is capped by the part of the workload that isn't accelerated.

How to use LFM2.5-VL-DSpark