Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data

Gintare Karolina Dziugaite, Daniel M. Roy
Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, PMLR R15:201-210, 2017.

Abstract

One of the defining properties of deep learn- ing is that models are chosen to have many more parameters than available training data. In light of this capacity for overfitting, it is remarkable that simple algorithms like SGD re- liably return solutions with low test error. One roadblock to explaining these phenomena in terms of implicit regularization, structural prop- erties of the solution, and/or easiness of the data is that many learning bounds are quan- titatively vacuous when applied to networks learned by SGD in this “deep learning” regime. Logically, in order to explain generalization, we need nonvacuous bounds. We return to an idea by Langford and Caruana (2001), who used PAC-Bayes bounds to compute nonvac- uous numerical bounds on generalization error for stochastic two-layer two-hidden-unit neural networks via a sensitivity analysis. By optimiz- ing the PAC-Bayes bound directly, we are able to extend their approach and obtain nonvacu- ous generalization bounds for deep stochastic neural network classifiers with millions of pa- rameters trained on only tens of thousands of examples. We connect our findings to recent and old work on flat minima and MDL-based explanations of generalization.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR15-dziugaite17a, title = {Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data}, author = {Dziugaite, Gintare Karolina and Roy, Daniel M.}, booktitle = {Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence}, pages = {201--210}, year = {2017}, editor = {Elidan, Gal and Kersting, Kristian}, volume = {R15}, series = {Proceedings of Machine Learning Research}, month = {11--15 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r15/main/assets/dziugaite17a/dziugaite17a.pdf}, url = {https://proceedings.mlr.press/r15/dziugaite17a.html}, abstract = {One of the defining properties of deep learn- ing is that models are chosen to have many more parameters than available training data. In light of this capacity for overfitting, it is remarkable that simple algorithms like SGD re- liably return solutions with low test error. One roadblock to explaining these phenomena in terms of implicit regularization, structural prop- erties of the solution, and/or easiness of the data is that many learning bounds are quan- titatively vacuous when applied to networks learned by SGD in this “deep learning” regime. Logically, in order to explain generalization, we need nonvacuous bounds. We return to an idea by Langford and Caruana (2001), who used PAC-Bayes bounds to compute nonvac- uous numerical bounds on generalization error for stochastic two-layer two-hidden-unit neural networks via a sensitivity analysis. By optimiz- ing the PAC-Bayes bound directly, we are able to extend their approach and obtain nonvacu- ous generalization bounds for deep stochastic neural network classifiers with millions of pa- rameters trained on only tens of thousands of examples. We connect our findings to recent and old work on flat minima and MDL-based explanations of generalization.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data %A Gintare Karolina Dziugaite %A Daniel M. Roy %B Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2017 %E Gal Elidan %E Kristian Kersting %F pmlr-vR15-dziugaite17a %I PMLR %P 201--210 %U https://proceedings.mlr.press/r15/dziugaite17a.html %V R15 %X One of the defining properties of deep learn- ing is that models are chosen to have many more parameters than available training data. In light of this capacity for overfitting, it is remarkable that simple algorithms like SGD re- liably return solutions with low test error. One roadblock to explaining these phenomena in terms of implicit regularization, structural prop- erties of the solution, and/or easiness of the data is that many learning bounds are quan- titatively vacuous when applied to networks learned by SGD in this “deep learning” regime. Logically, in order to explain generalization, we need nonvacuous bounds. We return to an idea by Langford and Caruana (2001), who used PAC-Bayes bounds to compute nonvac- uous numerical bounds on generalization error for stochastic two-layer two-hidden-unit neural networks via a sensitivity analysis. By optimiz- ing the PAC-Bayes bound directly, we are able to extend their approach and obtain nonvacu- ous generalization bounds for deep stochastic neural network classifiers with millions of pa- rameters trained on only tens of thousands of examples. We connect our findings to recent and old work on flat minima and MDL-based explanations of generalization. %Z Reissued by PMLR on 04 October 2026.
APA
Dziugaite, G.K. & Roy, D.M.. (2017). Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data. Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R15:201-210 Available from https://proceedings.mlr.press/r15/dziugaite17a.html. Reissued by PMLR on 04 October 2026.

Related Material