SGLang

2 articles tagged with SGLang

September 24, 2026
analysisLiquid Ai

Liquid AI Releases DSpark Draft Model for LFM2.5-VL-3B, Claims Up to 3.13x Decode Speedup

Liquid AI has released LFM2.5-VL-3B-DSpark, a 280M-parameter speculative decoding drafter for its LFM2.5-VL-3B vision-language model. The company claims decode speedups up to 3.13x on Apple silicon and 2.66x on H100 GPUs, with day-one support for llama.cpp, MLX-VLM, and SGLang.

August 20, 2026
changelogLiquid Ai

Liquid AI Ships DSpark Draft Models, Cutting LFM2.5 Inference Latency Up to 3.18x on GPU

Liquid AI released DSpark draft model checkpoints for three LFM2.5 models, enabling speculative decoding that speeds up inference by up to 3.18x on H100 GPUs and 2.87x on-device, with day-one support for llama.cpp and SGLang.