Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

Zirui Ge, Pengxiang Ding, Baohua Yin, Yemin Wang, Qishen Wang, Zhiyong Xie, Hengtao Li, Runze Suo, Wenxuan Song, Han Zhao, Shangke Lyu, Haoang Li, Ran Cheng, Cheng Chi, Hui-Bin Ge, Yaozhi Luo, Donglin Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:34393-34406, 2026.

Abstract

Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynamics needed for manipulation, so small prediction errors can be amplified by downstream policies. We propose Dyn-VPP, a post-training framework that casts multi-step denoising as policy optimization and aligns predicted future latents with expert visual dynamics via verifiable terminal reward, without modifying any architecture. This enables explicit optimization of dynamics signals that are not captured by likelihood-only training. As a result, Dyn-VPP yields more accurate visual dynamics and improves downstream task execution. Experiments across diverse simulated and real-world manipulation settings show improved dynamics consistency and consistently higher task success.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ge26a, title = {Dyn-{VPP}: Video Prediction Policy Optimization for Improved Visual Dynamics}, author = {Ge, Zirui and Ding, Pengxiang and Yin, Baohua and Wang, Yemin and Wang, Qishen and Xie, Zhiyong and Li, Hengtao and Suo, Runze and Song, Wenxuan and Zhao, Han and Lyu, Shangke and Li, Haoang and Cheng, Ran and Chi, Cheng and Ge, Hui-Bin and Luo, Yaozhi and Wang, Donglin}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {34393--34406}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ge26a/ge26a.pdf}, url = {https://proceedings.mlr.press/v306/ge26a.html}, abstract = {Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynamics needed for manipulation, so small prediction errors can be amplified by downstream policies. We propose Dyn-VPP, a post-training framework that casts multi-step denoising as policy optimization and aligns predicted future latents with expert visual dynamics via verifiable terminal reward, without modifying any architecture. This enables explicit optimization of dynamics signals that are not captured by likelihood-only training. As a result, Dyn-VPP yields more accurate visual dynamics and improves downstream task execution. Experiments across diverse simulated and real-world manipulation settings show improved dynamics consistency and consistently higher task success.} }
Endnote
%0 Conference Paper %T Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics %A Zirui Ge %A Pengxiang Ding %A Baohua Yin %A Yemin Wang %A Qishen Wang %A Zhiyong Xie %A Hengtao Li %A Runze Suo %A Wenxuan Song %A Han Zhao %A Shangke Lyu %A Haoang Li %A Ran Cheng %A Cheng Chi %A Hui-Bin Ge %A Yaozhi Luo %A Donglin Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ge26a %I PMLR %P 34393--34406 %U https://proceedings.mlr.press/v306/ge26a.html %V 306 %X Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynamics needed for manipulation, so small prediction errors can be amplified by downstream policies. We propose Dyn-VPP, a post-training framework that casts multi-step denoising as policy optimization and aligns predicted future latents with expert visual dynamics via verifiable terminal reward, without modifying any architecture. This enables explicit optimization of dynamics signals that are not captured by likelihood-only training. As a result, Dyn-VPP yields more accurate visual dynamics and improves downstream task execution. Experiments across diverse simulated and real-world manipulation settings show improved dynamics consistency and consistently higher task success.
APA
Ge, Z., Ding, P., Yin, B., Wang, Y., Wang, Q., Xie, Z., Li, H., Suo, R., Song, W., Zhao, H., Lyu, S., Li, H., Cheng, R., Chi, C., Ge, H., Luo, Y. & Wang, D.. (2026). Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:34393-34406 Available from https://proceedings.mlr.press/v306/ge26a.html.

Related Material