Cross-Embodiment Robot Foundation World Models with Latent Actions

Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Tsung-Yen Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, Franziska Meier
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:48175-48191, 2026.

Abstract

The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model’s performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results shows that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM’s downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlights the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-huang26bv, title = {Cross-Embodiment Robot Foundation World Models with Latent Actions}, author = {Huang, Huang and Yenamandra, Sriram and Majumdar, Arjun and Aljalbout, Elie and Nagarajan, Tushar and Yang, Tsung-Yen and Rai, Akshara and Rabbat, Michael and Fei-Fei, Li and Wu, Jiajun and Wu, Tingfan and Meier, Franziska}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {48175--48191}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/huang26bv/huang26bv.pdf}, url = {https://proceedings.mlr.press/v306/huang26bv.html}, abstract = {The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model’s performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results shows that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM’s downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlights the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics.} }
Endnote
%0 Conference Paper %T Cross-Embodiment Robot Foundation World Models with Latent Actions %A Huang Huang %A Sriram Yenamandra %A Arjun Majumdar %A Elie Aljalbout %A Tushar Nagarajan %A Tsung-Yen Yang %A Akshara Rai %A Michael Rabbat %A Li Fei-Fei %A Jiajun Wu %A Tingfan Wu %A Franziska Meier %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-huang26bv %I PMLR %P 48175--48191 %U https://proceedings.mlr.press/v306/huang26bv.html %V 306 %X The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model’s performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results shows that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM’s downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlights the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics.
APA
Huang, H., Yenamandra, S., Majumdar, A., Aljalbout, E., Nagarajan, T., Yang, T., Rai, A., Rabbat, M., Fei-Fei, L., Wu, J., Wu, T. & Meier, F.. (2026). Cross-Embodiment Robot Foundation World Models with Latent Actions. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:48175-48191 Available from https://proceedings.mlr.press/v306/huang26bv.html.

Related Material