Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning

Ming Chen, Sheng Tang, Rong-Xi Tan, Ziniu Li, Jiacheng Chen, Ke Xue, Chao Qian
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:16875-16909, 2026.

Abstract

Decoding-based regression, which reformulates regression as a sequence generation task, has emerged as a promising paradigm of applying large language models for numerical prediction. However, its progress is hindered by the misalignment between discrete token-level objectives (e.g., cross-entropy) and continuous numerical values. Existing approaches relying on token-level constraints often fail to capture the global magnitude of the target value, limiting their precision and generalization. In this paper, we propose to unlock the potential of decoding-based regression via reinforcement learning. We formulate the generation process as a Markov decision process, utilizing sequence-level rewards to enforce global numerical coherence.Under this framework, we present GenRe$^2$, which combines policy gradient methods and on-policy distillation to provide dense expert supervision while preserving error magnitudes, thereby resolving the temporal credit assignment challenge. Extensive experiments across tabular regression, code metric prediction and generative reward modeling demonstrate that GenRe$^2$ consistently outperforms traditional baselines, establishing a robust paradigm for general-purpose numerical prediction.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ec, title = {Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning}, author = {Chen, Ming and Tang, Sheng and Tan, Rong-Xi and Li, Ziniu and Chen, Jiacheng and Xue, Ke and Qian, Chao}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {16875--16909}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ec/chen26ec.pdf}, url = {https://proceedings.mlr.press/v306/chen26ec.html}, abstract = {Decoding-based regression, which reformulates regression as a sequence generation task, has emerged as a promising paradigm of applying large language models for numerical prediction. However, its progress is hindered by the misalignment between discrete token-level objectives (e.g., cross-entropy) and continuous numerical values. Existing approaches relying on token-level constraints often fail to capture the global magnitude of the target value, limiting their precision and generalization. In this paper, we propose to unlock the potential of decoding-based regression via reinforcement learning. We formulate the generation process as a Markov decision process, utilizing sequence-level rewards to enforce global numerical coherence.Under this framework, we present GenRe$^2$, which combines policy gradient methods and on-policy distillation to provide dense expert supervision while preserving error magnitudes, thereby resolving the temporal credit assignment challenge. Extensive experiments across tabular regression, code metric prediction and generative reward modeling demonstrate that GenRe$^2$ consistently outperforms traditional baselines, establishing a robust paradigm for general-purpose numerical prediction.} }
Endnote
%0 Conference Paper %T Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning %A Ming Chen %A Sheng Tang %A Rong-Xi Tan %A Ziniu Li %A Jiacheng Chen %A Ke Xue %A Chao Qian %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ec %I PMLR %P 16875--16909 %U https://proceedings.mlr.press/v306/chen26ec.html %V 306 %X Decoding-based regression, which reformulates regression as a sequence generation task, has emerged as a promising paradigm of applying large language models for numerical prediction. However, its progress is hindered by the misalignment between discrete token-level objectives (e.g., cross-entropy) and continuous numerical values. Existing approaches relying on token-level constraints often fail to capture the global magnitude of the target value, limiting their precision and generalization. In this paper, we propose to unlock the potential of decoding-based regression via reinforcement learning. We formulate the generation process as a Markov decision process, utilizing sequence-level rewards to enforce global numerical coherence.Under this framework, we present GenRe$^2$, which combines policy gradient methods and on-policy distillation to provide dense expert supervision while preserving error magnitudes, thereby resolving the temporal credit assignment challenge. Extensive experiments across tabular regression, code metric prediction and generative reward modeling demonstrate that GenRe$^2$ consistently outperforms traditional baselines, establishing a robust paradigm for general-purpose numerical prediction.
APA
Chen, M., Tang, S., Tan, R., Li, Z., Chen, J., Xue, K. & Qian, C.. (2026). Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:16875-16909 Available from https://proceedings.mlr.press/v306/chen26ec.html.

Related Material