
Models & tools
Accelerating vision-language models with LFM2.5-VL-DSpark
Publisher preview
The vision drafter uses the same architecture as our text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to draft a block of k candidate tokens. Image patches and text tokens are projected into a shared representation before those layers, so the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality. The inference algorithm is therefore unchanged from the text models. We follow the DSpark recipe with a mixture of vision-language SFT data, weighted toward…





