SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning

Xiaojun Guo, Runyu Zhou, Yifei Wang, Qi Zhang, Chenheng Zhang, Stefanie Jegelka, Xiaohan Wang, Jiajun Chai, Guojun Yin, Wei Lin, Yisen Wang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:38720-38745, 2026.

Abstract

Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often rely on textual shortcuts rather than adequately using visual evidence during reasoning. Although reinforcement learning (RL) can align models with desired behaviors, its application to VLMs has been hindered by the lack of scalable and reliable rewards. To overcome this challenge, we propose SSL4RL, a novel framework that leverages self-supervised learning (SSL) tasks as a source of verifiable rewards for RL. Our approach reformulates SSL objectives like rotation prediction and patch reconstruction into dense automatic rewards, removing the need for human preferences or AI evaluators. Experiments show that SSL4RL substantially improves performance on both vision-centric and vision-language reasoning benchmarks, with encouraging potentials on open-ended scenarios and stronger resilience to visual corruptions. Through systematic ablations, we identify key factors influencing SSL4RL, including data volume, model scale, model choice, task combination, and task difficulty, thereby offering new design principles for future work. Our implementation is open-sourced at https://github.com/PKU-ML/SSL4RL, with models hosted on Huggingface collection https://huggingface.co/collections/PKU-ML/ssl4rl.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-guo26ak, title = {{SSL}4{RL}: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning}, author = {Guo, Xiaojun and Zhou, Runyu and Wang, Yifei and Zhang, Qi and Zhang, Chenheng and Jegelka, Stefanie and Wang, Xiaohan and Chai, Jiajun and Yin, Guojun and Lin, Wei and Wang, Yisen}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {38720--38745}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/guo26ak/guo26ak.pdf}, url = {https://proceedings.mlr.press/v306/guo26ak.html}, abstract = {Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often rely on textual shortcuts rather than adequately using visual evidence during reasoning. Although reinforcement learning (RL) can align models with desired behaviors, its application to VLMs has been hindered by the lack of scalable and reliable rewards. To overcome this challenge, we propose SSL4RL, a novel framework that leverages self-supervised learning (SSL) tasks as a source of verifiable rewards for RL. Our approach reformulates SSL objectives like rotation prediction and patch reconstruction into dense automatic rewards, removing the need for human preferences or AI evaluators. Experiments show that SSL4RL substantially improves performance on both vision-centric and vision-language reasoning benchmarks, with encouraging potentials on open-ended scenarios and stronger resilience to visual corruptions. Through systematic ablations, we identify key factors influencing SSL4RL, including data volume, model scale, model choice, task combination, and task difficulty, thereby offering new design principles for future work. Our implementation is open-sourced at https://github.com/PKU-ML/SSL4RL, with models hosted on Huggingface collection https://huggingface.co/collections/PKU-ML/ssl4rl.} }
Endnote
%0 Conference Paper %T SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning %A Xiaojun Guo %A Runyu Zhou %A Yifei Wang %A Qi Zhang %A Chenheng Zhang %A Stefanie Jegelka %A Xiaohan Wang %A Jiajun Chai %A Guojun Yin %A Wei Lin %A Yisen Wang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-guo26ak %I PMLR %P 38720--38745 %U https://proceedings.mlr.press/v306/guo26ak.html %V 306 %X Vision-language models (VLMs) have shown remarkable abilities by integrating large language models with visual inputs. However, they often rely on textual shortcuts rather than adequately using visual evidence during reasoning. Although reinforcement learning (RL) can align models with desired behaviors, its application to VLMs has been hindered by the lack of scalable and reliable rewards. To overcome this challenge, we propose SSL4RL, a novel framework that leverages self-supervised learning (SSL) tasks as a source of verifiable rewards for RL. Our approach reformulates SSL objectives like rotation prediction and patch reconstruction into dense automatic rewards, removing the need for human preferences or AI evaluators. Experiments show that SSL4RL substantially improves performance on both vision-centric and vision-language reasoning benchmarks, with encouraging potentials on open-ended scenarios and stronger resilience to visual corruptions. Through systematic ablations, we identify key factors influencing SSL4RL, including data volume, model scale, model choice, task combination, and task difficulty, thereby offering new design principles for future work. Our implementation is open-sourced at https://github.com/PKU-ML/SSL4RL, with models hosted on Huggingface collection https://huggingface.co/collections/PKU-ML/ssl4rl.
APA
Guo, X., Zhou, R., Wang, Y., Zhang, Q., Zhang, C., Jegelka, S., Wang, X., Chai, J., Yin, G., Lin, W. & Wang, Y.. (2026). SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:38720-38745 Available from https://proceedings.mlr.press/v306/guo26ak.html.

Related Material