Representative, Informative, and De-Amplifying: Requirements for Robust Bayesian Active Learning under Model Misspecification

Roubing Tang, Sabina J. Sloman, Samuel Kaski
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3016-3024, 2026.

Abstract

In many science and industry settings, a central challenge is designing experiments under time and budget constraints. \emph{Bayesian Optimal Experimental Design (BOED)} is a paradigm to pick maximally informative designs that has been widely applied to such problems. During training, BOED selects inputs according to a pre-determined acquisition criterion to target \emph{informativeness}. During testing, the model learned during training encounters a naturally occurring distribution of test samples. This leads to an instance of covariate shift, where the train and test samples are drawn from different distributions (the training samples are not \emph{representative} of the test distribution). Prior work has shown that in the presence of model misspecification, covariate shift amplifies generalization error. Our first contribution is to provide a mathematical analysis of generalization error in the presence of model misspecification, revealing that, beyond covariate shift, generalization error is also driven by a previously unidentified phenomenon we term \emph{error (de-)amplification}. We then develop a new acquisition function that mitigates the effects of model misspecification by including terms for representativeness, informativeness, and de-amplification (R-IDeA). Our experimental results demonstrate that the proposed method performs better than methods that target only informativeness, only representativeness, or both.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-tang26d, title = { Representative, Informative, and De-Amplifying: Requirements for Robust Bayesian Active Learning under Model Misspecification }, author = {Tang, Roubing and Sloman, Sabina J. and Kaski, Samuel}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3016--3024}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/tang26d/tang26d.pdf}, url = {https://proceedings.mlr.press/v300/tang26d.html}, abstract = { In many science and industry settings, a central challenge is designing experiments under time and budget constraints. \emph{Bayesian Optimal Experimental Design (BOED)} is a paradigm to pick maximally informative designs that has been widely applied to such problems. During training, BOED selects inputs according to a pre-determined acquisition criterion to target \emph{informativeness}. During testing, the model learned during training encounters a naturally occurring distribution of test samples. This leads to an instance of covariate shift, where the train and test samples are drawn from different distributions (the training samples are not \emph{representative} of the test distribution). Prior work has shown that in the presence of model misspecification, covariate shift amplifies generalization error. Our first contribution is to provide a mathematical analysis of generalization error in the presence of model misspecification, revealing that, beyond covariate shift, generalization error is also driven by a previously unidentified phenomenon we term \emph{error (de-)amplification}. We then develop a new acquisition function that mitigates the effects of model misspecification by including terms for representativeness, informativeness, and de-amplification (R-IDeA). Our experimental results demonstrate that the proposed method performs better than methods that target only informativeness, only representativeness, or both. } }
Endnote
%0 Conference Paper %T Representative, Informative, and De-Amplifying: Requirements for Robust Bayesian Active Learning under Model Misspecification %A Roubing Tang %A Sabina J. Sloman %A Samuel Kaski %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-tang26d %I PMLR %P 3016--3024 %U https://proceedings.mlr.press/v300/tang26d.html %V 300 %X In many science and industry settings, a central challenge is designing experiments under time and budget constraints. \emph{Bayesian Optimal Experimental Design (BOED)} is a paradigm to pick maximally informative designs that has been widely applied to such problems. During training, BOED selects inputs according to a pre-determined acquisition criterion to target \emph{informativeness}. During testing, the model learned during training encounters a naturally occurring distribution of test samples. This leads to an instance of covariate shift, where the train and test samples are drawn from different distributions (the training samples are not \emph{representative} of the test distribution). Prior work has shown that in the presence of model misspecification, covariate shift amplifies generalization error. Our first contribution is to provide a mathematical analysis of generalization error in the presence of model misspecification, revealing that, beyond covariate shift, generalization error is also driven by a previously unidentified phenomenon we term \emph{error (de-)amplification}. We then develop a new acquisition function that mitigates the effects of model misspecification by including terms for representativeness, informativeness, and de-amplification (R-IDeA). Our experimental results demonstrate that the proposed method performs better than methods that target only informativeness, only representativeness, or both.
APA
Tang, R., Sloman, S.J. & Kaski, S.. (2026). Representative, Informative, and De-Amplifying: Requirements for Robust Bayesian Active Learning under Model Misspecification . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3016-3024 Available from https://proceedings.mlr.press/v300/tang26d.html.

Related Material