SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

Amir Dellali, Luca A Lanzendörfer, Florian Grötschla, Roger Wattenhofer
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:23653-23671, 2026.

Abstract

We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dellali26a, title = {{SALSA}-V: Shortcut-Augmented Long-form Synchronized Audio from Videos}, author = {Dellali, Amir and Lanzend\"{o}rfer, Luca A and Gr\"{o}tschla, Florian and Wattenhofer, Roger}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {23653--23671}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dellali26a/dellali26a.pdf}, url = {https://proceedings.mlr.press/v306/dellali26a.html}, abstract = {We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design.} }
Endnote
%0 Conference Paper %T SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos %A Amir Dellali %A Luca A Lanzendörfer %A Florian Grötschla %A Roger Wattenhofer %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dellali26a %I PMLR %P 23653--23671 %U https://proceedings.mlr.press/v306/dellali26a.html %V 306 %X We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design.
APA
Dellali, A., Lanzendörfer, L.A., Grötschla, F. & Wattenhofer, R.. (2026). SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:23653-23671 Available from https://proceedings.mlr.press/v306/dellali26a.html.

Related Material