Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

Mostafa Elhoushi, Alexander D. Pretko, Nolan Simran Dey, Bin Claire Zhang, Gavia Gray, Gurpreet Gosal, Abdulrahman Mahmoud, Shane Bergsma, Joel Hestness
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27844-27860, 2026.

Abstract

Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout—particularly layer dropout—has largely disappeared from LLM pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, LLM can achieve lower or similar validation loss while saving upto 20% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.4$\times$ inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-elhoushi26a, title = {Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient {LLM} Training and Inference}, author = {Elhoushi, Mostafa and Pretko, Alexander D. and Dey, Nolan Simran and Zhang, Bin Claire and Gray, Gavia and Gosal, Gurpreet and Mahmoud, Abdulrahman and Bergsma, Shane and Hestness, Joel}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27844--27860}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/elhoushi26a/elhoushi26a.pdf}, url = {https://proceedings.mlr.press/v306/elhoushi26a.html}, abstract = {Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout—particularly layer dropout—has largely disappeared from LLM pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, LLM can achieve lower or similar validation loss while saving upto 20% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.4$\times$ inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.} }
Endnote
%0 Conference Paper %T Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference %A Mostafa Elhoushi %A Alexander D. Pretko %A Nolan Simran Dey %A Bin Claire Zhang %A Gavia Gray %A Gurpreet Gosal %A Abdulrahman Mahmoud %A Shane Bergsma %A Joel Hestness %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-elhoushi26a %I PMLR %P 27844--27860 %U https://proceedings.mlr.press/v306/elhoushi26a.html %V 306 %X Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout—particularly layer dropout—has largely disappeared from LLM pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, LLM can achieve lower or similar validation loss while saving upto 20% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.4$\times$ inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.
APA
Elhoushi, M., Pretko, A.D., Dey, N.S., Zhang, B.C., Gray, G., Gosal, G., Mahmoud, A., Bergsma, S. & Hestness, J.. (2026). Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27844-27860 Available from https://proceedings.mlr.press/v306/elhoushi26a.html.

Related Material