VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

Hongyang Du, Junjie Ye, Xiaoyan Cong, Runhao Li, Jingcheng Ni, Aman Agarwal, Zeqi Zhou, Zekun Li, Randall Balestriero, Yue Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:26797-26832, 2026.

Abstract

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, geometric plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-du26q, title = {{V}ideo{GPA}: Distilling Geometry Priors for 3{D}-Consistent Video Generation}, author = {Du, Hongyang and Ye, Junjie and Cong, Xiaoyan and Li, Runhao and Ni, Jingcheng and Agarwal, Aman and Zhou, Zeqi and Li, Zekun and Balestriero, Randall and Wang, Yue}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {26797--26832}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/du26q/du26q.pdf}, url = {https://proceedings.mlr.press/v306/du26q.html}, abstract = {While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, geometric plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments.} }
Endnote
%0 Conference Paper %T VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation %A Hongyang Du %A Junjie Ye %A Xiaoyan Cong %A Runhao Li %A Jingcheng Ni %A Aman Agarwal %A Zeqi Zhou %A Zekun Li %A Randall Balestriero %A Yue Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-du26q %I PMLR %P 26797--26832 %U https://proceedings.mlr.press/v306/du26q.html %V 306 %X While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, geometric plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments.
APA
Du, H., Ye, J., Cong, X., Li, R., Ni, J., Agarwal, A., Zhou, Z., Li, Z., Balestriero, R. & Wang, Y.. (2026). VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:26797-26832 Available from https://proceedings.mlr.press/v306/du26q.html.

Related Material