SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling

Xiang Fang, Wanlong Fang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:29015-29028, 2026.

Abstract

In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental “temporal gap”, rendering models blind to critical causal transitions. Existing solutions relying on generative hallucination (e.g., latent diffusion) or autoregressive extrapolation often fail to maintain semantic consistency over long horizons, suffering from object vanishing and energetic instability. We propose a paradigm shift from probabilistic generation to variational mechanics with the Semantic Least Action Principle (SLAP). Drawing a rigorous isomorphism between classical mechanics and semantic dynamics, we model the latent video trajectory as a path on a Riemannian manifold governed by a Semantic Lagrangian. By formulating the interpolation task as a Boundary Value Problem (BVP) solved via the discrete Euler-Lagrange equations, SLAP naturally enforces object persistence without pixel-level rendering. Extensive experiments on multiple challenging datasets show the effectiveness of our proposed SLAP.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-fang26c, title = {{SLAP}: The Semantic Least Action Principle for Variational Video-Language Modeling}, author = {Fang, Xiang and Fang, Wanlong}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {29015--29028}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/fang26c/fang26c.pdf}, url = {https://proceedings.mlr.press/v306/fang26c.html}, abstract = {In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental “temporal gap”, rendering models blind to critical causal transitions. Existing solutions relying on generative hallucination (e.g., latent diffusion) or autoregressive extrapolation often fail to maintain semantic consistency over long horizons, suffering from object vanishing and energetic instability. We propose a paradigm shift from probabilistic generation to variational mechanics with the Semantic Least Action Principle (SLAP). Drawing a rigorous isomorphism between classical mechanics and semantic dynamics, we model the latent video trajectory as a path on a Riemannian manifold governed by a Semantic Lagrangian. By formulating the interpolation task as a Boundary Value Problem (BVP) solved via the discrete Euler-Lagrange equations, SLAP naturally enforces object persistence without pixel-level rendering. Extensive experiments on multiple challenging datasets show the effectiveness of our proposed SLAP.} }
Endnote
%0 Conference Paper %T SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling %A Xiang Fang %A Wanlong Fang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-fang26c %I PMLR %P 29015--29028 %U https://proceedings.mlr.press/v306/fang26c.html %V 306 %X In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental “temporal gap”, rendering models blind to critical causal transitions. Existing solutions relying on generative hallucination (e.g., latent diffusion) or autoregressive extrapolation often fail to maintain semantic consistency over long horizons, suffering from object vanishing and energetic instability. We propose a paradigm shift from probabilistic generation to variational mechanics with the Semantic Least Action Principle (SLAP). Drawing a rigorous isomorphism between classical mechanics and semantic dynamics, we model the latent video trajectory as a path on a Riemannian manifold governed by a Semantic Lagrangian. By formulating the interpolation task as a Boundary Value Problem (BVP) solved via the discrete Euler-Lagrange equations, SLAP naturally enforces object persistence without pixel-level rendering. Extensive experiments on multiple challenging datasets show the effectiveness of our proposed SLAP.
APA
Fang, X. & Fang, W.. (2026). SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:29015-29028 Available from https://proceedings.mlr.press/v306/fang26c.html.

Related Material