Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models

Wenhui Tan, Fiorenzo Parascandolo, Enver Sangineto, Jianzhong Ju, Zhenbo Luo, Qian Cao, Rita Cucchiara, Ruihua Song, Jian Luan
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:118588-118608, 2026.

Abstract

Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases pass@$n$ accuracy. Empirically, the final-layer posterior of post-trained LRMs exhibit sharply reduced entropy, while the entropy of intermediate layers remains relatively high. Motivated by this entropy asymmetry, we propose Latent Exploration Decoding (LED), a depth-conditioned decoding strategy. LED aggregates intermediate posteriors via cumulative sum and selects depth configurations with maximal entropy as exploration candidates. Without additional training or parameters, LED consistently improves pass@1 and pass@16 accuracy by 0.61 and 1.03 percentage points across multiple reasoning benchmarks and models. Relevant code is included in the supplementary material and will made be fully public after this paper is accepted.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-tan26d, title = {Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models}, author = {Tan, Wenhui and Parascandolo, Fiorenzo and Sangineto, Enver and Ju, Jianzhong and Luo, Zhenbo and Cao, Qian and Cucchiara, Rita and Song, Ruihua and Luan, Jian}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {118588--118608}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/tan26d/tan26d.pdf}, url = {https://proceedings.mlr.press/v306/tan26d.html}, abstract = {Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases pass@$n$ accuracy. Empirically, the final-layer posterior of post-trained LRMs exhibit sharply reduced entropy, while the entropy of intermediate layers remains relatively high. Motivated by this entropy asymmetry, we propose Latent Exploration Decoding (LED), a depth-conditioned decoding strategy. LED aggregates intermediate posteriors via cumulative sum and selects depth configurations with maximal entropy as exploration candidates. Without additional training or parameters, LED consistently improves pass@1 and pass@16 accuracy by 0.61 and 1.03 percentage points across multiple reasoning benchmarks and models. Relevant code is included in the supplementary material and will made be fully public after this paper is accepted.} }
Endnote
%0 Conference Paper %T Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models %A Wenhui Tan %A Fiorenzo Parascandolo %A Enver Sangineto %A Jianzhong Ju %A Zhenbo Luo %A Qian Cao %A Rita Cucchiara %A Ruihua Song %A Jian Luan %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-tan26d %I PMLR %P 118588--118608 %U https://proceedings.mlr.press/v306/tan26d.html %V 306 %X Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases pass@$n$ accuracy. Empirically, the final-layer posterior of post-trained LRMs exhibit sharply reduced entropy, while the entropy of intermediate layers remains relatively high. Motivated by this entropy asymmetry, we propose Latent Exploration Decoding (LED), a depth-conditioned decoding strategy. LED aggregates intermediate posteriors via cumulative sum and selects depth configurations with maximal entropy as exploration candidates. Without additional training or parameters, LED consistently improves pass@1 and pass@16 accuracy by 0.61 and 1.03 percentage points across multiple reasoning benchmarks and models. Relevant code is included in the supplementary material and will made be fully public after this paper is accepted.
APA
Tan, W., Parascandolo, F., Sangineto, E., Ju, J., Luo, Z., Cao, Q., Cucchiara, R., Song, R. & Luan, J.. (2026). Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:118588-118608 Available from https://proceedings.mlr.press/v306/tan26d.html.

Related Material