Combining predictions from linear models when training and test inputs differ

Thijs Van Ommen CWI Amsterdam
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:912-921, 2014.

Abstract

Methods for combining predictions from dif- ferent models in a supervised learning setting must somehow estimate/predict the quality of a model’s predictions at unknown future inputs. Many of these methods (often implicitly) make the assumption that the test inputs are identical to the training inputs, which is seldom reasonable. By failing to take into account that prediction will generally be harder for test inputs that did not occur in the training set, this leads to the se- lection of too complex models. Based on a novel, unbiased expression for KL divergence, we pro- pose XAIC and its special case FAIC as versions of AIC intended for prediction that use different degrees of knowledge of the test inputs. Both methods substantially differ from and may out- perform all the known versions of AIC even when the training and test inputs are iid, and are es- pecially useful for deterministic inputs and under covariate shift. Our experiments on linear models suggest that if the test and training inputs differ substantially, then XAIC and FAIC predictively outperform AIC, BIC and several other methods including Bayesian model averaging.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-amsterdam14a, title = {Combining predictions from linear models when training and test inputs differ}, author = {Amsterdam, Thijs Van Ommen CWI}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {912--921}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/amsterdam14a/amsterdam14a.pdf}, url = {https://proceedings.mlr.press/r12/amsterdam14a.html}, abstract = {Methods for combining predictions from dif- ferent models in a supervised learning setting must somehow estimate/predict the quality of a model’s predictions at unknown future inputs. Many of these methods (often implicitly) make the assumption that the test inputs are identical to the training inputs, which is seldom reasonable. By failing to take into account that prediction will generally be harder for test inputs that did not occur in the training set, this leads to the se- lection of too complex models. Based on a novel, unbiased expression for KL divergence, we pro- pose XAIC and its special case FAIC as versions of AIC intended for prediction that use different degrees of knowledge of the test inputs. Both methods substantially differ from and may out- perform all the known versions of AIC even when the training and test inputs are iid, and are es- pecially useful for deterministic inputs and under covariate shift. Our experiments on linear models suggest that if the test and training inputs differ substantially, then XAIC and FAIC predictively outperform AIC, BIC and several other methods including Bayesian model averaging.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Combining predictions from linear models when training and test inputs differ %A Thijs Van Ommen CWI Amsterdam %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-amsterdam14a %I PMLR %P 912--921 %U https://proceedings.mlr.press/r12/amsterdam14a.html %V R12 %X Methods for combining predictions from dif- ferent models in a supervised learning setting must somehow estimate/predict the quality of a model’s predictions at unknown future inputs. Many of these methods (often implicitly) make the assumption that the test inputs are identical to the training inputs, which is seldom reasonable. By failing to take into account that prediction will generally be harder for test inputs that did not occur in the training set, this leads to the se- lection of too complex models. Based on a novel, unbiased expression for KL divergence, we pro- pose XAIC and its special case FAIC as versions of AIC intended for prediction that use different degrees of knowledge of the test inputs. Both methods substantially differ from and may out- perform all the known versions of AIC even when the training and test inputs are iid, and are es- pecially useful for deterministic inputs and under covariate shift. Our experiments on linear models suggest that if the test and training inputs differ substantially, then XAIC and FAIC predictively outperform AIC, BIC and several other methods including Bayesian model averaging. %Z Reissued by PMLR on 04 October 2026.
APA
Amsterdam, T.V.O.C.. (2014). Combining predictions from linear models when training and test inputs differ. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:912-921 Available from https://proceedings.mlr.press/r12/amsterdam14a.html. Reissued by PMLR on 04 October 2026.

Related Material