Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic Scheduling

Jingchu Gai, Guanning Zeng, Huaqing Zhang, Han Zhong, Yige Hong, Andrej Risteski, Aditi Raghunathan
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:32614-32639, 2026.

Abstract

This paper investigates a pivotal yet debated component of reinforcement learning (RL) for training large language models (LLMs): controlling entropy (increasing or decreasing it) during RL fine-tuning. The existing literature presents a dichotomy: some studies posit that increasing entropy facilitates exploration, whereas others argue that decreasing entropy enhances performance. To reconcile these conflicting observations, we provide a theoretical framework showing that the effect of entropy is governed by Entropy Discrepancy, the distributional divergence between positive and negative samples. Guided by this insight, we derive a principled dynamic scheduling method that adaptively modulates the entropy coefficient, effectively switching between entropy maximization and minimization as training evolves. Extensive experiments confirm the correlation between Entropy Discrepancy and the efficacy of entropy control. Furthermore, our adaptive method yields substantial improvements, boosting Pass@K by 6.7% on AIME24 and 17.52% on puzzle tasks compared to vanilla RL, while consistently outperforming recent state-of-the-art reasoning methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gai26b, title = {Demystifying Entropy Control in {LLM} {RL} Training: Theoretical Analysis and Dynamic Scheduling}, author = {Gai, Jingchu and Zeng, Guanning and Zhang, Huaqing and Zhong, Han and Hong, Yige and Risteski, Andrej and Raghunathan, Aditi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {32614--32639}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gai26b/gai26b.pdf}, url = {https://proceedings.mlr.press/v306/gai26b.html}, abstract = {This paper investigates a pivotal yet debated component of reinforcement learning (RL) for training large language models (LLMs): controlling entropy (increasing or decreasing it) during RL fine-tuning. The existing literature presents a dichotomy: some studies posit that increasing entropy facilitates exploration, whereas others argue that decreasing entropy enhances performance. To reconcile these conflicting observations, we provide a theoretical framework showing that the effect of entropy is governed by Entropy Discrepancy, the distributional divergence between positive and negative samples. Guided by this insight, we derive a principled dynamic scheduling method that adaptively modulates the entropy coefficient, effectively switching between entropy maximization and minimization as training evolves. Extensive experiments confirm the correlation between Entropy Discrepancy and the efficacy of entropy control. Furthermore, our adaptive method yields substantial improvements, boosting Pass@K by 6.7% on AIME24 and 17.52% on puzzle tasks compared to vanilla RL, while consistently outperforming recent state-of-the-art reasoning methods.} }
Endnote
%0 Conference Paper %T Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic Scheduling %A Jingchu Gai %A Guanning Zeng %A Huaqing Zhang %A Han Zhong %A Yige Hong %A Andrej Risteski %A Aditi Raghunathan %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gai26b %I PMLR %P 32614--32639 %U https://proceedings.mlr.press/v306/gai26b.html %V 306 %X This paper investigates a pivotal yet debated component of reinforcement learning (RL) for training large language models (LLMs): controlling entropy (increasing or decreasing it) during RL fine-tuning. The existing literature presents a dichotomy: some studies posit that increasing entropy facilitates exploration, whereas others argue that decreasing entropy enhances performance. To reconcile these conflicting observations, we provide a theoretical framework showing that the effect of entropy is governed by Entropy Discrepancy, the distributional divergence between positive and negative samples. Guided by this insight, we derive a principled dynamic scheduling method that adaptively modulates the entropy coefficient, effectively switching between entropy maximization and minimization as training evolves. Extensive experiments confirm the correlation between Entropy Discrepancy and the efficacy of entropy control. Furthermore, our adaptive method yields substantial improvements, boosting Pass@K by 6.7% on AIME24 and 17.52% on puzzle tasks compared to vanilla RL, while consistently outperforming recent state-of-the-art reasoning methods.
APA
Gai, J., Zeng, G., Zhang, H., Zhong, H., Hong, Y., Risteski, A. & Raghunathan, A.. (2026). Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic Scheduling. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:32614-32639 Available from https://proceedings.mlr.press/v306/gai26b.html.

Related Material