DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

Silin Gao, Hao Zhao, Zeming Chen, Sepideh Mamooler, Antara Raaghavi Bhattacharya, Qiyu Wu, Hiromi Wakaki, Yuki Mitsufuji, Li Mi, Syrielle Montariol, Antoine Bosselut
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:33953-33984, 2026.

Abstract

Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW’s superior visual dynamic understanding boosts its downstream performances on both visual narrative creation and world simulation, showing improved consistency and controllability of visual generation and better instruction-following ability.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gao26am, title = {{D}yna{V}ie{W}: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics}, author = {Gao, Silin and Zhao, Hao and Chen, Zeming and Mamooler, Sepideh and Bhattacharya, Antara Raaghavi and Wu, Qiyu and Wakaki, Hiromi and Mitsufuji, Yuki and Mi, Li and Montariol, Syrielle and Bosselut, Antoine}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {33953--33984}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gao26am/gao26am.pdf}, url = {https://proceedings.mlr.press/v306/gao26am.html}, abstract = {Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW’s superior visual dynamic understanding boosts its downstream performances on both visual narrative creation and world simulation, showing improved consistency and controllability of visual generation and better instruction-following ability.} }
Endnote
%0 Conference Paper %T DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics %A Silin Gao %A Hao Zhao %A Zeming Chen %A Sepideh Mamooler %A Antara Raaghavi Bhattacharya %A Qiyu Wu %A Hiromi Wakaki %A Yuki Mitsufuji %A Li Mi %A Syrielle Montariol %A Antoine Bosselut %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gao26am %I PMLR %P 33953--33984 %U https://proceedings.mlr.press/v306/gao26am.html %V 306 %X Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require models to predict or simulate multiple levels of dynamic constituents, such as actions taken in the visual sequence, and the associated changes to the visual environment that result. To address this challenge, we propose a dynamic schema-guided world model, DynaVieW, optimized for visual dynamic prediction and simulation. DynaVieW achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes, and transitions capture comprehensive dynamic constituents within a hierarchical schema. DynaVieW jointly models transition prediction and state simulation under a mixture-of-experts architecture, with a cross-expert selective attention and a schema token re-weighted loss, to ensure effective and robust learning. DynaVieW’s superior visual dynamic understanding boosts its downstream performances on both visual narrative creation and world simulation, showing improved consistency and controllability of visual generation and better instruction-following ability.
APA
Gao, S., Zhao, H., Chen, Z., Mamooler, S., Bhattacharya, A.R., Wu, Q., Wakaki, H., Mitsufuji, Y., Mi, L., Montariol, S. & Bosselut, A.. (2026). DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:33953-33984 Available from https://proceedings.mlr.press/v306/gao26am.html.

Related Material