RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning

Yukun Chen, Jiaming Li, Longze Chen, Ze Gong, Jingpeng Li, Zhen Qin, Hengyu Chang, Lei Zhang, Ancheng Xu, Zhihao Yang, Hamid Alinejad-Rokny, Qiang Qu, Bo Zheng, Min Yang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:15060-15090, 2026.

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a prevailing paradigm for enhancing reasoning in Multimodal Large Language Models (MLLMs). However, relying solely on outcome supervision risks reward hacking, where models learn spurious reasoning patterns to satisfy final answer checks. While recent rubric-based approaches offer fine-grained supervision signals, they suffer from high computational costs of instance-level generation and inefficient training dynamics caused by treating all rubrics as equally learnable. In this paper, we propose Stratified Rubric-based Curriculum Learning (RuCL), a novel framework that reformulates curriculum learning by shifting the focus from data selection to reward design. RuCL generates generalized rubrics for broad applicability and stratifies them based on model competence, dynamically adjusting their weights to guide the model from foundational perception to advanced logical reasoning. Extensive experiments on various visual reasoning benchmarks show that RuCL yields a remarkable +7.83% average improvement over the Qwen2.5-VL-7B model, achieving a state-of-the-art accuracy of 60.06%.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26bk, title = {{R}u{CL}: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning}, author = {Chen, Yukun and Li, Jiaming and Chen, Longze and Gong, Ze and Li, Jingpeng and Qin, Zhen and Chang, Hengyu and Zhang, Lei and Xu, Ancheng and Yang, Zhihao and Alinejad-Rokny, Hamid and Qu, Qiang and Zheng, Bo and Yang, Min}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {15060--15090}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26bk/chen26bk.pdf}, url = {https://proceedings.mlr.press/v306/chen26bk.html}, abstract = {Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a prevailing paradigm for enhancing reasoning in Multimodal Large Language Models (MLLMs). However, relying solely on outcome supervision risks reward hacking, where models learn spurious reasoning patterns to satisfy final answer checks. While recent rubric-based approaches offer fine-grained supervision signals, they suffer from high computational costs of instance-level generation and inefficient training dynamics caused by treating all rubrics as equally learnable. In this paper, we propose Stratified Rubric-based Curriculum Learning (RuCL), a novel framework that reformulates curriculum learning by shifting the focus from data selection to reward design. RuCL generates generalized rubrics for broad applicability and stratifies them based on model competence, dynamically adjusting their weights to guide the model from foundational perception to advanced logical reasoning. Extensive experiments on various visual reasoning benchmarks show that RuCL yields a remarkable +7.83% average improvement over the Qwen2.5-VL-7B model, achieving a state-of-the-art accuracy of 60.06%.} }
Endnote
%0 Conference Paper %T RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning %A Yukun Chen %A Jiaming Li %A Longze Chen %A Ze Gong %A Jingpeng Li %A Zhen Qin %A Hengyu Chang %A Lei Zhang %A Ancheng Xu %A Zhihao Yang %A Hamid Alinejad-Rokny %A Qiang Qu %A Bo Zheng %A Min Yang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26bk %I PMLR %P 15060--15090 %U https://proceedings.mlr.press/v306/chen26bk.html %V 306 %X Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a prevailing paradigm for enhancing reasoning in Multimodal Large Language Models (MLLMs). However, relying solely on outcome supervision risks reward hacking, where models learn spurious reasoning patterns to satisfy final answer checks. While recent rubric-based approaches offer fine-grained supervision signals, they suffer from high computational costs of instance-level generation and inefficient training dynamics caused by treating all rubrics as equally learnable. In this paper, we propose Stratified Rubric-based Curriculum Learning (RuCL), a novel framework that reformulates curriculum learning by shifting the focus from data selection to reward design. RuCL generates generalized rubrics for broad applicability and stratifies them based on model competence, dynamically adjusting their weights to guide the model from foundational perception to advanced logical reasoning. Extensive experiments on various visual reasoning benchmarks show that RuCL yields a remarkable +7.83% average improvement over the Qwen2.5-VL-7B model, achieving a state-of-the-art accuracy of 60.06%.
APA
Chen, Y., Li, J., Chen, L., Gong, Z., Li, J., Qin, Z., Chang, H., Zhang, L., Xu, A., Yang, Z., Alinejad-Rokny, H., Qu, Q., Zheng, B. & Yang, M.. (2026). RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:15060-15090 Available from https://proceedings.mlr.press/v306/chen26bk.html.

Related Material