AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Liquid AI released an experimental DSpark speculative-decoding drafter for its LFM2.5-VL-3B vision-language model. The 280M-parameter add-on speeds decoding up to 3.13x on Apple silicon and 2.66x on H100 without changing outputs, with day-one llama.cpp, MLX-VLM, and SGLang support.

Liquid AI has released LFM2.5-VL-DSpark, an experimental draft model that accelerates inference of its open-weight LFM2.5-VL-3B vision-language model through speculative decoding. According to the company, the 280M-parameter drafter adds just 8.9% to the target model’s parameter count while delivering decoding speedups of up to 3.13x on device and up to 2.66x on an NVIDIA H100 — without changing output quality. The model is available now on Hugging Face in Safetensors and GGUF formats.

The drafter extends Liquid AI’s DSpark recipe — previously applied to its text-only LFM2.5 models — to a multimodal target. It captures the target model’s hidden states at a fixed set of tapped layers and drafts blocks of candidate tokens conditioned on those states. Because image patches and text tokens are projected into a shared representation before the tapped layers, the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality, and the inference algorithm is unchanged from the text models.

Trained on a mixture of vision-language supervised fine-tuning data weighted toward expected serving workloads, the final drafter is a simplified attention-only model with 4 layers, selected via ablations across 3, 4, and 5 layers, with a block size of 9. Liquid AI reports that acceptance improved over 10 training epochs before reaching diminishing returns. The drafter comprises a 193.0M-parameter decoder stack, a 21.0M hidden-state projection, a 65.5M Markov head, and roughly 6.4k parameters in norms and a confidence head — about 279.5M in total.

Benchmarks follow the MMSpec protocol across six vision tasks: general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation. On-device, with MLX on an M5 Max, decoding runs 2.30x to 3.13x faster and end-to-end latency improves 1.56x to 2.62x. With llama.cpp on an M3 Ultra, decoding improves 1.57x to 2.14x, end-to-end 1.30x to 1.77x. On H100, the company reports decoding speedups ranging up to 2.66x and end-to-end gains of 1.64x to 2.27x, using a DSpark block size of 8. Speculative decoding is exact: the target model verifies every proposed token, so greedy output matches the target alone.

At a glance
announcementWhen: announced September 2026, available now
The developmentLiquid AI announced and released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, available immediately on Hugging Face.
Accelerating Vision-Language Models With LFM2.5-VL-DSpark
Speculative Decoding · Liquid AI · September 2026

Accelerating Vision-Language Models With LFM2.5-VL-DSpark

Liquid AI releases an experimental DSpark drafter for its open-weight LFM2.5-VL-3B vision-language model. The 280M-parameter add-on speeds decoding up to 3.13× on Apple silicon and 2.66× on H100 — without changing outputs — with day-one llama.cpp, MLX-VLM, and SGLang support.

3.13×
Peak decode speedup — MLX on M5 Max
2.66×
Peak decode speedup — NVIDIA H100
+8.9%
Added to target model’s parameter count
280M
Drafter Params
4
Attention Layers
9
Draft Block Size
10
Training Epochs
6
Benchmark Tasks
01 — Mechanism

How Speculative Decoding Buys Speed

A small, fast draft model proposes candidate tokens that the larger target model verifies in batch — accepting matches and discarding the rest. Because verification is cheaper than sequential generation, accepted drafts yield net speedup with identical outputs under greedy decoding.

1

Tap Hidden States

DSpark captures the target model’s hidden states at a fixed set of tapped layers.

2

Project & Draft

Image patches and text tokens share one representation, so the drafter handles either modality in identical dimensionality.

3

Block Proposal

The drafter proposes blocks of candidate tokens (block size 8–9) in a single pass.

4

Exact Verification

The target verifies every proposed token — greedy output matches the target alone.

02 — Architecture

A Simplified Attention-Only Drafter

Layer count was chosen via ablations across 3, 4, and 5 layers; acceptance improved over 10 training epochs before hitting diminishing returns. Trained on vision-language SFT data weighted toward expected serving workloads.

Decoder Stack

Attention-only, 4 layers
193.0M

Hidden-State Projection

Maps tapped target states into drafter space
21.0M

Markov Head

Generates block-structured candidate tokens
65.5M

Norms + Confidence Head

Roughly 6.4k parameters — total ≈ 279.5M
~6.4k
03 — Benchmarks (MMSpec Protocol)

Reported Speedups Across Six Vision Tasks

Evaluated on general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation. Figures are Liquid AI’s own measurements, not independently verified.

Decode · MLX / M5 Max
2.30–3.13×
End-to-End · MLX / M5 Max
1.56–2.62×
Decode · llama.cpp / M3 Ultra
1.57–2.14×
End-to-End · llama.cpp / M3 Ultra
1.30–1.77×
Decode · H100 · Block 8
up to 2.66×
End-to-End · H100 · Block 8
1.64–2.27×

Note: end-to-end gains trail decode gains because speculative decoding accelerates only the decode phase — not vision encoding or prefill. Liquid AI itself highlights this Amdahl’s-law ceiling.

04 — Why It Matters

Faster Vision AI on Consumer Hardware

Vision-language models run slower than text models because images pass through a vision encoder and are processed as hundreds of visual tokens. A drafter that doubles or triples decode speed — for under 9% more parameters — could make 3B-class multimodal models practical on laptops and phones.

Ecosystem

Day-One Integration

DSpark support ships in llama.cpp (PR #29339), MLX-VLM (PR #2280), and SGLang (PR #40651) — acceleration works in toolchains hobbyists and deployers already use, no custom inference code required.

Licensing

Open-Weight, Unrestricted

Download, fine-tune, and deploy without restrictions, per Liquid AI. Available now on Hugging Face in both Safetensors and GGUF formats, with OpenAI-compatible endpoints via SGLang, llama-server, and mlx_vlm.server.

Recipe

From Text to Multimodal

DSpark drafting was previously applied to text-only LFM2.5 models. The vision drafter reuses the same architecture and inference algorithm, differing mainly in training data and the shared multimodal representation.

05 — In Their Words

From the Announcement

“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”

— Liquid AI, announcement post

“Speculative decoding speeds up only decode, not vision encoding or prefill. When those stages already take up much of the wall time, even a large decode speedup gives only a modest end-to-end gain.”

— Liquid AI, announcement post

“With LFM2.5, we’re delivering on our vision of AI that runs anywhere.”

— Liquid AI, announcement post
06 — Assessment

Strengths vs. Open Questions

The release is explicitly labeled experimental, with no stated timeline for a stable promotion. Here is how the evidence breaks down.

What Supports the Claims

  • Mechanism is well understood: exact speculative decoding with a small drafter is a proven speed-for-memory trade
  • Company candidly frames its own Amdahl’s-law limits
  • Day-one llama.cpp, MLX-VLM, and SGLang integration makes it practically usable, not paper-only
  • Open-weight licensing with unrestricted fine-tuning and deployment

Caveats & Gaps

  • Self-reported benchmarks on hardware most readers don’t own — no third-party verification
  • An H100 decode range printed as “20.4x to 2.66x” reads as a typo (likely 2.04×); unclarified by Liquid AI
  • Acceptance rates on out-of-distribution vision tasks unknown
  • Effect on sampling-based (non-greedy) generation quality and runtime memory footprint unpublished; training data mixture not disclosed
07 — Release at a Glance

Key Facts

Dimension Detail Status
Announced September 2026 — available immediately on Hugging Face Live now
Formats Safetensors and GGUF; OpenAI-compatible endpoints via SGLang, llama-server, mlx_vlm.server Shipped
Recommended block sizes 8 or 9, depending on hardware Guidance
Stability Explicitly experimental; no timeline for stable promotion; recipe may still change Experimental
What to watch Independent consumer-hardware benchmarks, H100 figure clarification, extension to other LFM2.5 sizes and modalities Pending

Faster Vision AI on Consumer Hardware

The release targets a persistent bottleneck for local and edge AI: vision-language models are slower than text models because images must pass through a vision encoder and then be processed as hundreds of visual tokens alongside the text prompt. A drafter that roughly doubles or triples decode speed — for under 9% more parameters — could make 3B-class multimodal models practical on laptops and phones, where users feel latency directly.

Day-one integration matters as much as the speedup itself. DSpark support in llama.cpp, MLX-VLM, and SGLang means the acceleration works in the toolchains hobbyists and deployers already use, rather than requiring custom inference code. Combined with Liquid AI’s open-weight licensing — download, fine-tune, and deploy without restrictions, per the company — the release positions the LFM2.5 family for on-device multimodal applications.

The company is also candid about the ceiling: speculative decoding accelerates only the decode phase, not vision encoding or prefill. On edge devices with limited compute, those stages consume a larger share of wall-clock time, so end-to-end gains (1.30x–2.62x) trail decode gains — an illustration of Amdahl’s law that Liquid AI itself highlights.

Amazon

Top picks for "accelerat vision language"

As an affiliate, we earn on qualifying purchases.

From Text Drafters to Multimodal

: “

Speculative decoding is an established technique in which a small, fast “draft” model proposes candidate tokens that the larger target model verifies in batch, accepting matching tokens and discarding the rest. Because verification is cheaper than sequential generation, accepted drafts translate into net speedup with identical outputs under greedy decoding.

Liquid AI released its first LFM2.5-DSpark drafter models for text-only LFM2.5 targets earlier in 2026. The new vision drafter reuses the same architecture and inference algorithm, differing mainly in training data and the shared multimodal representation. The LFM2.5 family spans base models, audio, and vision variants, with the 3B vision model positioned as an edge-capable multimodal option.

“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”

— Liquid AI, announcement post

Experimental Status and Benchmark Gaps

The release is explicitly labeled experimental, and the company has not indicated when or whether the drafter will be promoted to a stable release. The reported speedups are Liquid AI’s own measurements, not independently verified third-party benchmarks, and results will vary with hardware, prompt composition, and image resolution.

One figure in the company’s GPU results appears internally inconsistent — a lower bound described as “20.4x” in a range stated as “20.4x to 2.66x” — which reads as a typo, likely for 2.04x; Liquid AI has not clarified the figure. It is also unclear how acceptance rates behave on out-of-distribution vision tasks, how the drafter affects sampling-based (non-greedy) generation quality, and what the memory footprint increase is in runtime terms beyond parameter count. Details of the training data mixture have not been published.

Community Uptake and Clarifications

The model and required integration patches are public: SGLang support requires a build with DSpark for LFM2 targets (PR #40651), llama.cpp requires PR #29339, and MLX-VLM requires PR #2280. Early adopters will likely publish independent benchmarks on consumer hardware, which will test whether Liquid AI’s reported ranges hold outside the company’s own setup.

Watch for clarification of the inconsistent H100 decode figure, broader hardware coverage, and whether Liquid AI extends DSpark to other model sizes or modalities in the LFM2.5 family. Users can run the models today via the OpenAI-compatible endpoints exposed by SGLang, llama-server, and mlx_vlm.server, with block sizes of 8 or 9 recommended depending on hardware.

Where I land

I find this release credible in its core claim because the mechanism is well understood: exact speculative decoding with a small drafter is a proven way to buy speed with minimal memory, and Liquid AI’s own framing of Amdahl’s-law limits lends the announcement honesty that hype-driven releases often lack. The day-one integration across llama.cpp, MLX-VLM, and SGLang is the detail that makes this practically useful rather than a paper-only result.

The strongest counterargument is that these are self-reported benchmarks on hardware most readers don’t own (an M5 Max, an M3 Ultra, an H100), with an internally inconsistent H100 figure and no third-party verification. Real-world acceptance rates on diverse images and long prompts could fall well below the headline ranges, and the ‘experimental’ label signals the recipe may still change.

What would change my assessment: independent benchmarks reproducing the reported speedups on common consumer hardware, clarification of the H100 lower-bound figure, and published acceptance-rate data across a broader task distribution. If acceptance holds up outside Liquid AI’s curated mixture, this drafter becomes an easy default for anyone running LFM2.5-VL-3B locally.

Key Questions

Does the DSpark drafter change the model’s answers?

No, according to Liquid AI. Speculative decoding is exact: the target model verifies every proposed token, so greedy output is identical to running the target model alone. The company notes per-response timings report draft_n / draft_n_accepted so users can see acceptance behavior.

How much extra memory does the drafter need?

The drafter adds approximately 280M parameters, an 8.9% increase on the 3B target model. Liquid AI describes this as a minimal memory footprint trade for the speedup.

Which runtimes support it on day one?

llama.cpp (PR #29339), MLX-VLM (PR #2280), and SGLang (PR #40651) all ship LFM-compatible DSpark support. The model is downloadable from Hugging Face in Safetensors and GGUF formats.

Why are end-to-end gains smaller than decode speedups?

Speculative decoding only accelerates the decode phase. Vision encoding and prefill — which are proportionally more expensive on edge devices — remain unaccelerated, capping the overall speedup in line with Amdahl’s law.

What block size should I use?

Liquid AI recommends a block size of 8 or 9 depending on hardware. The draft model was trained with a block size of 9; runtimes read the value from the model’s config or sidecar metadata.

Source: Hugging Face

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

ChatGPT Ads Expands Across Europe

OpenAI has announced a European expansion of ChatGPT Ads, but countries, formats, privacy rules and rollout dates remain unconfirmed.

Court Rules Pentagon Can Blacklist Anthropic For Refusing To Enable Claude Features – Ars Technica

An Ars Technica headline reports that a court ruled the Pentagon can blacklist Anthropic over Claude features. The ruling’s scope and reasoning remain unclear.

Anthropic’s Opus 4.6 Is A Smut-machine – TechCrunch

A TechCrunch report finds Anthropic’s Claude Opus 4.6 generates sexually explicit material with minimal prompting, testing the company’s safety stance.

How to Choose AI Automation Software For Small Businesses

Step-by-step guide to choosing, connecting, and launching AI automation software for your small business in 2-4 weeks.