Learning How Deep to Go: Self-Scaling Deep Reinforcement Learning

Michelangelo Vegliò, Marco Fantozzi, Antonio Di Cecco, Carlo Metta, Flora Angileri, Simone Treccani, Adrienne Chloe Rayos Macazar, Silvia Giulia Galfre’, Maurizio Parton, Francesco Morandin
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2809-2817, 2026.

Abstract

Deep Reinforcement Learning (DRL) has achieved remarkable results in complex sequential decision-making tasks, often using very deep neural networks. However, these architectures incur substantial computational and energy costs, and selecting the optimal network depth in advance remains an open challenge. In this paper, we introduce SCALE-RL, a self-scaling DRL framework that dynamically adjusts its architectural depth during training, allowing the network to automatically adapt its depth to the task. Integrated into an AlphaZero-style pipeline for Othello, our approach matches the playing strength of the baseline agent while reducing network depth by 50%. This process not only translates into substantial savings in computation and energy but also enhances model interpretability through the additive decomposition of decision-making across layers. Our results suggest that enabling DRL models to discover the complexity they require, rather than relying on fixed, over-parameterized architectures, makes it possible to develop more efficient, interpretable, and sustainable DRL agents.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-veglio26a, title = { Learning How Deep to Go: Self-Scaling Deep Reinforcement Learning }, author = {Vegli\`{o}, Michelangelo and Fantozzi, Marco and Di Cecco, Antonio and Metta, Carlo and Angileri, Flora and Treccani, Simone and Macazar, Adrienne Chloe Rayos and Galfre', Silvia Giulia and Parton, Maurizio and Morandin, Francesco}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2809--2817}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/veglio26a/veglio26a.pdf}, url = {https://proceedings.mlr.press/v300/veglio26a.html}, abstract = { Deep Reinforcement Learning (DRL) has achieved remarkable results in complex sequential decision-making tasks, often using very deep neural networks. However, these architectures incur substantial computational and energy costs, and selecting the optimal network depth in advance remains an open challenge. In this paper, we introduce SCALE-RL, a self-scaling DRL framework that dynamically adjusts its architectural depth during training, allowing the network to automatically adapt its depth to the task. Integrated into an AlphaZero-style pipeline for Othello, our approach matches the playing strength of the baseline agent while reducing network depth by 50%. This process not only translates into substantial savings in computation and energy but also enhances model interpretability through the additive decomposition of decision-making across layers. Our results suggest that enabling DRL models to discover the complexity they require, rather than relying on fixed, over-parameterized architectures, makes it possible to develop more efficient, interpretable, and sustainable DRL agents. } }
Endnote
%0 Conference Paper %T Learning How Deep to Go: Self-Scaling Deep Reinforcement Learning %A Michelangelo Vegliò %A Marco Fantozzi %A Antonio Di Cecco %A Carlo Metta %A Flora Angileri %A Simone Treccani %A Adrienne Chloe Rayos Macazar %A Silvia Giulia Galfre’ %A Maurizio Parton %A Francesco Morandin %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-veglio26a %I PMLR %P 2809--2817 %U https://proceedings.mlr.press/v300/veglio26a.html %V 300 %X Deep Reinforcement Learning (DRL) has achieved remarkable results in complex sequential decision-making tasks, often using very deep neural networks. However, these architectures incur substantial computational and energy costs, and selecting the optimal network depth in advance remains an open challenge. In this paper, we introduce SCALE-RL, a self-scaling DRL framework that dynamically adjusts its architectural depth during training, allowing the network to automatically adapt its depth to the task. Integrated into an AlphaZero-style pipeline for Othello, our approach matches the playing strength of the baseline agent while reducing network depth by 50%. This process not only translates into substantial savings in computation and energy but also enhances model interpretability through the additive decomposition of decision-making across layers. Our results suggest that enabling DRL models to discover the complexity they require, rather than relying on fixed, over-parameterized architectures, makes it possible to develop more efficient, interpretable, and sustainable DRL agents.
APA
Vegliò, M., Fantozzi, M., Di Cecco, A., Metta, C., Angileri, F., Treccani, S., Macazar, A.C.R., Galfre’, S.G., Parton, M. & Morandin, F.. (2026). Learning How Deep to Go: Self-Scaling Deep Reinforcement Learning . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2809-2817 Available from https://proceedings.mlr.press/v300/veglio26a.html.

Related Material