SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation

Wei Chen, Xingyu Guo, Shuang Li, Fuwei Zhang, Meng Yuan, Jing Fan, Zhao Zhang, Deqing Wang, Fuzhen Zhuang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:14377-14391, 2026.

Abstract

Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals to provide richer token-level evidence for generation. However, existing approaches largely rely on alignment-centric fusion and underexplore synergistic information across modalities. In practice, synergistic information plays a critical role in capturing emergent item properties that cannot be inferred from any single modality alone. Such properties encode intrinsic item semantics and guide user preferences, enabling models to move beyond surface-level feature matching. To address this limitation, we propose SynGR, a synergistic generative recommendation framework that explicitly encourages the exploitation of cross-modal dependencies during generation. By constraining overreliance on dominant modalities, SynGR enables the model to capture emergent item semantics beyond shared or modality-specific signals. Extensive experiments across three benchmark datasets demonstrate that SynGR achieves superior performance.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ak, title = {{S}yn{GR}: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation}, author = {Chen, Wei and Guo, Xingyu and Li, Shuang and Zhang, Fuwei and Yuan, Meng and Fan, Jing and Zhang, Zhao and Wang, Deqing and Zhuang, Fuzhen}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {14377--14391}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ak/chen26ak.pdf}, url = {https://proceedings.mlr.press/v306/chen26ak.html}, abstract = {Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals to provide richer token-level evidence for generation. However, existing approaches largely rely on alignment-centric fusion and underexplore synergistic information across modalities. In practice, synergistic information plays a critical role in capturing emergent item properties that cannot be inferred from any single modality alone. Such properties encode intrinsic item semantics and guide user preferences, enabling models to move beyond surface-level feature matching. To address this limitation, we propose SynGR, a synergistic generative recommendation framework that explicitly encourages the exploitation of cross-modal dependencies during generation. By constraining overreliance on dominant modalities, SynGR enables the model to capture emergent item semantics beyond shared or modality-specific signals. Extensive experiments across three benchmark datasets demonstrate that SynGR achieves superior performance.} }
Endnote
%0 Conference Paper %T SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation %A Wei Chen %A Xingyu Guo %A Shuang Li %A Fuwei Zhang %A Meng Yuan %A Jing Fan %A Zhao Zhang %A Deqing Wang %A Fuzhen Zhuang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ak %I PMLR %P 14377--14391 %U https://proceedings.mlr.press/v306/chen26ak.html %V 306 %X Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals to provide richer token-level evidence for generation. However, existing approaches largely rely on alignment-centric fusion and underexplore synergistic information across modalities. In practice, synergistic information plays a critical role in capturing emergent item properties that cannot be inferred from any single modality alone. Such properties encode intrinsic item semantics and guide user preferences, enabling models to move beyond surface-level feature matching. To address this limitation, we propose SynGR, a synergistic generative recommendation framework that explicitly encourages the exploitation of cross-modal dependencies during generation. By constraining overreliance on dominant modalities, SynGR enables the model to capture emergent item semantics beyond shared or modality-specific signals. Extensive experiments across three benchmark datasets demonstrate that SynGR achieves superior performance.
APA
Chen, W., Guo, X., Li, S., Zhang, F., Yuan, M., Fan, J., Zhang, Z., Wang, D. & Zhuang, F.. (2026). SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:14377-14391 Available from https://proceedings.mlr.press/v306/chen26ak.html.

Related Material