Variance reduction combining pre-experiment and in-experiment data

Zhexiao Lin, Pablo Crespo
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:699-717, 2026.

Abstract

Online controlled experiments (A/B testing) are fundamental to data-driven decision-making in many companies. Improving the sensitivity of these experiments under fixed sample size constraints requires reducing the variance of the average treatment effect (ATE) estimator. Existing variance reduction techniques such as CUPED and CUPAC use pre-experiment data, but their effectiveness depends on how predictive those data are for outcomes measured during the experiment. In-experiment data are often more strongly correlated with the outcome, but using arbitrary post-treatment variables can introduce bias. In this paper, we propose a general, robust, and scalable framework that combines both pre-experiment and in-experiment data to achieve variance reduction. Our framework is simple, interpretable, and computationally efficient, making it practical for real-world deployment. We develop the asymptotic theory of the proposed estimator and provide consistent variance estimators. Empirical results from multiple online experiments conducted at Etsy demonstrate substantial additional variance reduction over current pipeline, even when incorporating only a few post-treatment covariates. These findings underscore the effectiveness of our framework in improving experimental sensitivity and accelerating data-driven decision-making.

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-lin26a, title = {Variance reduction combining pre-experiment and in-experiment data}, author = {Lin, Zhexiao and Crespo, Pablo}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {699--717}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/lin26a/lin26a.pdf}, url = {https://proceedings.mlr.press/v323/lin26a.html}, abstract = {Online controlled experiments (A/B testing) are fundamental to data-driven decision-making in many companies. Improving the sensitivity of these experiments under fixed sample size constraints requires reducing the variance of the average treatment effect (ATE) estimator. Existing variance reduction techniques such as CUPED and CUPAC use pre-experiment data, but their effectiveness depends on how predictive those data are for outcomes measured during the experiment. In-experiment data are often more strongly correlated with the outcome, but using arbitrary post-treatment variables can introduce bias. In this paper, we propose a general, robust, and scalable framework that combines both pre-experiment and in-experiment data to achieve variance reduction. Our framework is simple, interpretable, and computationally efficient, making it practical for real-world deployment. We develop the asymptotic theory of the proposed estimator and provide consistent variance estimators. Empirical results from multiple online experiments conducted at Etsy demonstrate substantial additional variance reduction over current pipeline, even when incorporating only a few post-treatment covariates. These findings underscore the effectiveness of our framework in improving experimental sensitivity and accelerating data-driven decision-making.} }
Endnote
%0 Conference Paper %T Variance reduction combining pre-experiment and in-experiment data %A Zhexiao Lin %A Pablo Crespo %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-lin26a %I PMLR %P 699--717 %U https://proceedings.mlr.press/v323/lin26a.html %V 323 %X Online controlled experiments (A/B testing) are fundamental to data-driven decision-making in many companies. Improving the sensitivity of these experiments under fixed sample size constraints requires reducing the variance of the average treatment effect (ATE) estimator. Existing variance reduction techniques such as CUPED and CUPAC use pre-experiment data, but their effectiveness depends on how predictive those data are for outcomes measured during the experiment. In-experiment data are often more strongly correlated with the outcome, but using arbitrary post-treatment variables can introduce bias. In this paper, we propose a general, robust, and scalable framework that combines both pre-experiment and in-experiment data to achieve variance reduction. Our framework is simple, interpretable, and computationally efficient, making it practical for real-world deployment. We develop the asymptotic theory of the proposed estimator and provide consistent variance estimators. Empirical results from multiple online experiments conducted at Etsy demonstrate substantial additional variance reduction over current pipeline, even when incorporating only a few post-treatment covariates. These findings underscore the effectiveness of our framework in improving experimental sensitivity and accelerating data-driven decision-making.
APA
Lin, Z. & Crespo, P.. (2026). Variance reduction combining pre-experiment and in-experiment data. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:699-717 Available from https://proceedings.mlr.press/v323/lin26a.html.

Related Material