WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

Weilun Feng, Guoxin Fan, Haotong Qin, Mingqiang Wu, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Dingrui Wang, Longlong Liao, Michele Magno, Yongjun Xu, Chuanguang Yang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:30040-30065, 2026.

Abstract

Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffusion transfer poorly to world models due to two world-model-specific obstacles: token heterogeneity from multi-modal coupling and spatial variation, and non-uniform temporal dynamics where a small set of hard tokens drives error growth, making uniform skipping either unstable or overly conservative. We propose WorldCache, a caching framework tailored to diffusion world models. We introduce Curvature-guided Heterogeneous Token Prediction, which uses a physics-grounded curvature score to estimate token predictability and applies a Hermite-guided damped predictor for chaotic tokens with abrupt direction changes. We also design Chaotic-prioritized Adaptive Skipping, which accumulates a curvature-normalized, dimensionless drift signal and recomputes only when bottleneck tokens begin to drift. Experiments on diffusion world models show that WorldCache delivers up to 3.7$\times$ end-to-end speedups while maintaining 98% rollout quality, demonstrating the vast advantages and practicality of WorldCache in resource-constrained scenarios.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-feng26d, title = {{W}orld{C}ache: Accelerating World Models for Free via Heterogeneous Token Caching}, author = {Feng, Weilun and Fan, Guoxin and Qin, Haotong and Wu, Mingqiang and Li, Yuqi and Li, Xiangqi and An, Zhulin and Huang, Libo and Wang, Dingrui and Liao, Longlong and Magno, Michele and Xu, Yongjun and Yang, Chuanguang}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {30040--30065}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/feng26d/feng26d.pdf}, url = {https://proceedings.mlr.press/v306/feng26d.html}, abstract = {Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffusion transfer poorly to world models due to two world-model-specific obstacles: token heterogeneity from multi-modal coupling and spatial variation, and non-uniform temporal dynamics where a small set of hard tokens drives error growth, making uniform skipping either unstable or overly conservative. We propose WorldCache, a caching framework tailored to diffusion world models. We introduce Curvature-guided Heterogeneous Token Prediction, which uses a physics-grounded curvature score to estimate token predictability and applies a Hermite-guided damped predictor for chaotic tokens with abrupt direction changes. We also design Chaotic-prioritized Adaptive Skipping, which accumulates a curvature-normalized, dimensionless drift signal and recomputes only when bottleneck tokens begin to drift. Experiments on diffusion world models show that WorldCache delivers up to 3.7$\times$ end-to-end speedups while maintaining 98% rollout quality, demonstrating the vast advantages and practicality of WorldCache in resource-constrained scenarios.} }
Endnote
%0 Conference Paper %T WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching %A Weilun Feng %A Guoxin Fan %A Haotong Qin %A Mingqiang Wu %A Yuqi Li %A Xiangqi Li %A Zhulin An %A Libo Huang %A Dingrui Wang %A Longlong Liao %A Michele Magno %A Yongjun Xu %A Chuanguang Yang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-feng26d %I PMLR %P 30040--30065 %U https://proceedings.mlr.press/v306/feng26d.html %V 306 %X Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffusion transfer poorly to world models due to two world-model-specific obstacles: token heterogeneity from multi-modal coupling and spatial variation, and non-uniform temporal dynamics where a small set of hard tokens drives error growth, making uniform skipping either unstable or overly conservative. We propose WorldCache, a caching framework tailored to diffusion world models. We introduce Curvature-guided Heterogeneous Token Prediction, which uses a physics-grounded curvature score to estimate token predictability and applies a Hermite-guided damped predictor for chaotic tokens with abrupt direction changes. We also design Chaotic-prioritized Adaptive Skipping, which accumulates a curvature-normalized, dimensionless drift signal and recomputes only when bottleneck tokens begin to drift. Experiments on diffusion world models show that WorldCache delivers up to 3.7$\times$ end-to-end speedups while maintaining 98% rollout quality, demonstrating the vast advantages and practicality of WorldCache in resource-constrained scenarios.
APA
Feng, W., Fan, G., Qin, H., Wu, M., Li, Y., Li, X., An, Z., Huang, L., Wang, D., Liao, L., Magno, M., Xu, Y. & Yang, C.. (2026). WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:30040-30065 Available from https://proceedings.mlr.press/v306/feng26d.html.

Related Material