Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

Xiaokun Feng, Jiashu Zhu, Meiqi Wu, Chubin Chen, Fangyuan Mao, Haiyang Guo, Jiahong Wu, Xiangxiang Chu, Kaiqi Huang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:30664-30691, 2026.

Abstract

Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memory consumption. However, the mismatch between training and inference, coupled with the challenge of maintaining long-term consistency, limits the effective utilization of foundation models. To mitigate these concerns, we propose MIGA, a novel infinite-frame long video generation method. Firstly, we propose an effective two-stage alignment mechanism that mitigates the training-inference gap by reducing the excessive noise span fed to the model. We then introduce an innovative dual consistency enhancement mechanism, where the self-reflection approach corrects early high-noise frames and the long-range frame guidance approach leverages later low-noise frames with broad coverage to steer generation, jointly improving temporal consistency. Extensive experiments on VBench and NarrLV demonstrate the state-of-the-art performance of MIGA. Our project page is available at https://xiaokunfeng.github.io/miga_homepage/.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-feng26ac, title = {Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos}, author = {Feng, Xiaokun and Zhu, Jiashu and Wu, Meiqi and Chen, Chubin and Mao, Fangyuan and Guo, Haiyang and Wu, Jiahong and Chu, Xiangxiang and Huang, Kaiqi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {30664--30691}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/feng26ac/feng26ac.pdf}, url = {https://proceedings.mlr.press/v306/feng26ac.html}, abstract = {Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memory consumption. However, the mismatch between training and inference, coupled with the challenge of maintaining long-term consistency, limits the effective utilization of foundation models. To mitigate these concerns, we propose MIGA, a novel infinite-frame long video generation method. Firstly, we propose an effective two-stage alignment mechanism that mitigates the training-inference gap by reducing the excessive noise span fed to the model. We then introduce an innovative dual consistency enhancement mechanism, where the self-reflection approach corrects early high-noise frames and the long-range frame guidance approach leverages later low-noise frames with broad coverage to steer generation, jointly improving temporal consistency. Extensive experiments on VBench and NarrLV demonstrate the state-of-the-art performance of MIGA. Our project page is available at https://xiaokunfeng.github.io/miga_homepage/.} }
Endnote
%0 Conference Paper %T Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos %A Xiaokun Feng %A Jiashu Zhu %A Meiqi Wu %A Chubin Chen %A Fangyuan Mao %A Haiyang Guo %A Jiahong Wu %A Xiangxiang Chu %A Kaiqi Huang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-feng26ac %I PMLR %P 30664--30691 %U https://proceedings.mlr.press/v306/feng26ac.html %V 306 %X Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memory consumption. However, the mismatch between training and inference, coupled with the challenge of maintaining long-term consistency, limits the effective utilization of foundation models. To mitigate these concerns, we propose MIGA, a novel infinite-frame long video generation method. Firstly, we propose an effective two-stage alignment mechanism that mitigates the training-inference gap by reducing the excessive noise span fed to the model. We then introduce an innovative dual consistency enhancement mechanism, where the self-reflection approach corrects early high-noise frames and the long-range frame guidance approach leverages later low-noise frames with broad coverage to steer generation, jointly improving temporal consistency. Extensive experiments on VBench and NarrLV demonstrate the state-of-the-art performance of MIGA. Our project page is available at https://xiaokunfeng.github.io/miga_homepage/.
APA
Feng, X., Zhu, J., Wu, M., Chen, C., Mao, F., Guo, H., Wu, J., Chu, X. & Huang, K.. (2026). Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:30664-30691 Available from https://proceedings.mlr.press/v306/feng26ac.html.

Related Material